{"id":128651,"url":"https://github.com/ai-boost/awesome-harness-engineering","name":"awesome-harness-engineering","description":"Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.","projects_count":468,"last_synced_at":"2026-09-08T10:00:23.129Z","repository":{"id":347812421,"uuid":"1195371847","full_name":"ai-boost/awesome-harness-engineering","owner":"ai-boost","description":"Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.","archived":false,"fork":false,"pushed_at":"2026-08-13T17:43:06.000Z","size":795,"stargazers_count":3545,"open_issues_count":148,"forks_count":421,"subscribers_count":28,"default_branch":"main","last_synced_at":"2026-08-13T21:56:50.509Z","etag":null,"topics":["agent-harness","agent-memory","agent-orchestration","ai-agent-harness","ai-agents","awesome-list","context-engineering","harness-engineering","mcp"],"latest_commit_sha":null,"homepage":"https://github.com/ai-boost/awesome-harness-engineering","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ai-boost.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","claude":"CLAUDE.md","gemini":null,"cursor":null,"copilot":null,"dco":null,"cla":null,"disclosure":null}},"created_at":"2026-03-29T15:39:49.000Z","updated_at":"2026-08-13T21:19:46.000Z","dependencies_parsed_at":"2026-08-13T19:27:18.888Z","dependency_job_id":null,"html_url":"https://github.com/ai-boost/awesome-harness-engineering","commit_stats":null,"previous_names":["ai-boost/awesome-harness-engineering"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/ai-boost/awesome-harness-engineering","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ai-boost%2Fawesome-harness-engineering","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ai-boost%2Fawesome-harness-engineering/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ai-boost%2Fawesome-harness-engineering/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ai-boost%2Fawesome-harness-engineering/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ai-boost","download_url":"https://codeload.github.com/ai-boost/awesome-harness-engineering/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ai-boost%2Fawesome-harness-engineering/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":341189360,"owners_count":37155667,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-08T02:00:05.531Z","response_time":108,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2026-04-18T20:30:34.443Z","updated_at":"2026-09-08T10:00:23.129Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Design Primitives","Related Awesome Lists","Security, Sandbox \u0026 Permissions","Evals \u0026 Verification","Reference Implementations","Foundations","Production Infrastructure \u0026 Operations","Acknowledgments"],"sub_categories":["Observability \u0026 Tracing","Planning \u0026 Task Decomposition","Agent Loop","Tool Design","Adjacent Collections","Verification \u0026 CI Integration","Memory \u0026 State","Task Runners \u0026 Orchestration","Skills \u0026 MCP","Context Delivery \u0026 Compaction","Debugging \u0026 Developer Experience","Demo Harnesses","Tutorials \u0026 Educational","Generators \u0026 Meta-Harnesses","Human-in-the-Loop","Permissions \u0026 Authorization"],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"assets/banner.jpg\" alt=\"Awesome Harness Engineering\" width=\"720\"\u003e\n  \u003ch1\u003eAwesome Harness Engineering\u003c/h1\u003e\n  \u003cp\u003eCurated resources, patterns, and templates for building reliable AI agent harnesses.\u003c/p\u003e\n  \u003cp\u003e\n    \u003ca href=\"https://awesome.re\"\u003e\u003cimg src=\"https://awesome.re/badge.svg\" alt=\"Awesome\"\u003e\u003c/a\u003e\n    \u003ca href=\"LICENSE\"\u003e\u003cimg src=\"https://img.shields.io/badge/License-CC0-lightgrey.svg\" alt=\"License: CC0\"\u003e\u003c/a\u003e\n    \u003ca href=\"https://github.com/ai-boost/awesome-harness-engineering/stargazers\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/ai-boost/awesome-harness-engineering?style=social\" alt=\"GitHub Stars\"\u003e\u003c/a\u003e\n    \u003ca href=\"https://github.com/ai-boost/awesome-harness-engineering/network/members\"\u003e\u003cimg src=\"https://img.shields.io/github/forks/ai-boost/awesome-harness-engineering?style=social\" alt=\"GitHub Forks\"\u003e\u003c/a\u003e\n    \u003ca href=\"https://github.com/ai-boost/awesome-harness-engineering/commits/main\"\u003e\u003cimg src=\"https://img.shields.io/github/last-commit/ai-boost/awesome-harness-engineering\" alt=\"Last Commit\"\u003e\u003c/a\u003e\n    \u003ca href=\"https://linux.do\"\u003e\u003cimg src=\"https://img.shields.io/badge/Join-linux.do-orange\" alt=\"linux.do\"\u003e\u003c/a\u003e\n  \u003c/p\u003e\n  \u003cp\u003e\n    \u003ca href=\"https://zdoc.app/de/ai-boost/awesome-harness-engineering\"\u003eDeutsch\u003c/a\u003e |\n    \u003ca href=\"https://zdoc.app/en/ai-boost/awesome-harness-engineering\"\u003eEnglish\u003c/a\u003e |\n    \u003ca href=\"https://zdoc.app/es/ai-boost/awesome-harness-engineering\"\u003eEspañol\u003c/a\u003e |\n    \u003ca href=\"https://zdoc.app/fr/ai-boost/awesome-harness-engineering\"\u003eFrançais\u003c/a\u003e |\n    \u003ca href=\"https://zdoc.app/ja/ai-boost/awesome-harness-engineering\"\u003e日本語\u003c/a\u003e |\n    \u003ca href=\"https://zdoc.app/ko/ai-boost/awesome-harness-engineering\"\u003e한국어\u003c/a\u003e |\n    \u003ca href=\"https://zdoc.app/pt/ai-boost/awesome-harness-engineering\"\u003ePortuguês\u003c/a\u003e |\n    \u003ca href=\"https://zdoc.app/ru/ai-boost/awesome-harness-engineering\"\u003eРусский\u003c/a\u003e |\n    \u003ca href=\"https://zdoc.app/zh/ai-boost/awesome-harness-engineering\"\u003e中文\u003c/a\u003e\n  \u003c/p\u003e\n\u003c/div\u003e\n\n**Harness engineering** is the discipline of designing the scaffolding — context delivery, tool interfaces, planning artifacts, verification loops, memory systems, and sandboxes — that surrounds an AI agent and determines whether it succeeds or fails on real tasks.\n\nThis list focuses on the *harness*, not the model. Every component here exists because the model can't do it alone — and the best harnesses are designed knowing those components will become unnecessary as models improve.\n\n---\n\n## Contents\n\n- [📐 Foundations](#foundations)\n- [🧩 Design Primitives](#design-primitives)\n  - [🔄 Agent Loop](#agent-loop)\n  - [🗺️ Planning \u0026 Task Decomposition](#planning--task-decomposition)\n  - [📦 Context Delivery \u0026 Compaction](#context-delivery--compaction)\n  - [🔧 Tool Design](#tool-design)\n  - [🔌 Skills \u0026 MCP](#skills--mcp)\n  - [🛡️ Permissions \u0026 Authorization](#permissions--authorization)\n  - [🧠 Memory \u0026 State](#memory--state)\n  - [⚙️ Task Runners \u0026 Orchestration](#task-runners--orchestration)\n  - [✔️ Verification \u0026 CI Integration](#verification--ci-integration)\n  - [👁️ Observability \u0026 Tracing](#observability--tracing)\n  - [🐛 Debugging \u0026 Developer Experience](#debugging--developer-experience)\n  - [🧑‍💼 Human-in-the-Loop](#human-in-the-loop)\n- [🔍 Reference Implementations](#reference-implementations)\n  - [🎓 Tutorials \u0026 Educational](#tutorials--educational)\n  - [🏭 Generators \u0026 Meta-Harnesses](#generators--meta-harnesses)\n  - [🧪 Demo Harnesses](#demo-harnesses)\n  - [🗂️ Adjacent Collections](#adjacent-collections)\n- [🔒 Security, Sandbox \u0026 Permissions](#security-sandbox--permissions)\n- [✅ Evals \u0026 Verification](#evals--verification)\n- [📋 Templates](#templates)\n- [📚 Related Awesome Lists](#related-awesome-lists)\n- [🤝 Contributing](#contributing)\n\n---\n\n## Foundations\n\nCanonical essays that define what harness engineering is and why it matters.\n\n- [Harness Engineering](https://openai.com/index/harness-engineering/) — OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.\n- [Unrolling the Codex Agent Loop](https://openai.com/index/unrolling-the-codex-agent-loop/) — OpenAI's detailed breakdown of the Codex agent loop, exposing each harness component and where it can be improved.\n- [Run Long-Horizon Tasks with Codex](https://developers.openai.com/blog/run-long-horizon-tasks-with-codex/) — OpenAI's practice guide for long-horizon task planning: introduces Plan.md, Implement.md, Documentation.md as reusable harness artifacts.\n- [Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) — Anthropic's foundational guide on agent architecture, covering when to use workflows vs. agents and how to compose primitives.\n- [Harness Design for Long-Running Application Development](https://www.anthropic.com/engineering/harness-design-long-running-apps) — Anthropic's engineering blog on designing harnesses for sustained, multi-session development tasks. Key insight: every harness component assumes the model can't do something; those assumptions expire.\n- [Writing Effective Tools for Agents](https://www.anthropic.com/engineering/writing-effective-tools-for-agents) — Anthropic's guide on tool interface design: naming, schemas, error surfaces, and the principle that tool design is agent UX.\n- [Beyond Permission Prompts](https://www.anthropic.com/engineering/beyond-permission-prompts) — Anthropic on building structured permission and authorization systems into agent harnesses instead of relying on natural-language permission text.\n- [Demystifying Evals for AI Agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) — Anthropic's framework for evaluating agent behavior: what to measure, how to build eval harnesses, and why unit-test-style evals fail for agents.\n- [What is an AI Agent?](https://www.ibm.com/think/topics/ai-agents) — IBM's definitional piece, useful for anchoring harness design decisions to a clear model of what an agent actually is.\n- [Agent Development Kit: Making it easy to build multi-agent applications](https://developers.googleblog.com/en/agent-development-kit-easy-to-build-multi-agent-applications/) — Google's announcement and design rationale for ADK: explains the multi-agent topology, tool registration model, and eval pipeline that shaped their framework. Complements the Anthropic/OpenAI framing with Google's production perspective.\n- [Harness Engineering](https://martinfowler.com/articles/exploring-gen-ai/harness-engineering.html) — Martin Fowler's synthesis of what harness engineering practice looks like: three interlocking systems — context engineering (curating what the agent knows), architectural constraints (deterministic linters and structural tests), and entropy management (periodic agents that repair documentation drift). The \"humans on the loop\" framing — harness engineers who design and maintain agent environments rather than inspecting individual outputs — is the clearest conceptual map of what the discipline actually entails.\n- [The Anatomy of an Agent Harness](https://blog.langchain.com/the-anatomy-of-an-agent-harness/) — LangChain's structural breakdown of the five primitives that compose a harness: filesystem (durable state + agent collaboration surface), code execution (autonomous problem-solving without pre-designed solutions), sandbox (isolation + verification), memory (cross-session persistence), and context management (compaction against \"context rot\"). The co-evolution warning — models trained with specific harnesses can become overfitted to those designs — explains why harness architecture choices have lasting consequences beyond the immediate task.\n- [Building AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned](https://arxiv.org/abs/2603.05344) — The first systematic practitioner paper on terminal-native coding agent harness design: eager-construction scaffolding (pre-build all components before the first message to eliminate first-call latency and race conditions), compound multi-model architecture (different model instances for execution, reasoning, critique, and vision tasks), 5-layer defense-in-depth safety, and schema-filtered planning subagents (enforce behavioral constraints via tool schema rather than runtime permission checks). The five lessons distilled from building OpenDev apply to any server-side agent harness.\n- [Natural-Language Agent Harnesses](https://arxiv.org/abs/2603.25723) — Proposes externalizing agent control logic as portable natural-language artifacts (NLAHs) executed by a shared Intelligent Harness Runtime, enabling harness design to be studied, transferred, and reproduced rather than buried in bespoke controller code. Directly addresses the root cause of harness fragility: control logic scattered across framework defaults and hard-coded controller logic that can't be inspected, versioned, or transferred.\n- [Ranking Engineer Agent (REA): Meta's Autonomous AI System for Ads Ranking](https://engineering.fb.com/2026/03/17/developer-tools/ranking-engineer-agent-rea-autonomous-ai-system-accelerating-meta-ads-ranking-innovation/) — Meta's production harness for multi-day ML pipeline automation with hibernate-and-wake checkpointing for resuming interrupted 6-hour tasks without losing context. Demonstrates harness design for scientific workflows where individual turns can exceed model context limits but the overall pipeline must maintain coherence across days.\n- [Supercharge Your AI Agents: The New ADK Integrations Ecosystem](https://developers.googleblog.com/en/supercharge-your-ai-agents-adk-integrations-ecosystem/) — Google's 2026 update to Agent Development Kit expanding the ecosystem integrations (Hugging Face, GitHub, Daytona, Notion, etc.) and providing reference patterns for how orchestration harnesses wire external services without losing determinism or state coherence.\n- [2026 Agentic Coding Trends Report](https://resources.anthropic.com/hubfs/2026%20Agentic%20Coding%20Trends%20Report.pdf?hsLang=en) — Anthropic's industry benchmark identifying infrastructure configuration as a first-class optimization variable: harness setup alone can swing benchmarks by 5+ percentage points. Documents the shift from single-agent to orchestrated multi-agent teams and introduces the \"agentic engineering platform\" category, bridging the gap between agent frameworks and production deployment infrastructure.\n- [How We Build Azure SRE Agent with Agentic Workflows](https://techcommunity.microsoft.com/blog/appsonazureblog/how-we-build-azure-sre-agent-with-agentic-workflows/4508753) — Architecture walkthrough of Microsoft's agent that has handled 35,000+ production incidents autonomously, reducing Azure App Service time-to-mitigation from 40.5 hours to 3 minutes. Documents the integration of MCP tools, telemetry, code repositories, and incident management platforms into a single agent harness with human-in-the-loop governance. The most data-backed production harness case study published in 2026.\n- [Context Engineering for Reliable AI Agents: Lessons from Building Azure SRE Agent](https://techcommunity.microsoft.com/blog/appsonazureblog/context-engineering-lessons-from-building-azure-sre-agent/4481200/) — Microsoft's account of shifting from 100+ bespoke tools and a prescriptive prompt to a filesystem-based context engineering system for their SRE agent. Key finding: exposing everything (source code, runbooks, query schemas, past investigation notes) as files and letting the agent use `read_file`, `grep`, `find`, and `shell` outperformed specialized tooling — \"Intent Met\" score rose from 45% to 75% on novel incidents.\n- [Harness Engineering: Structured Workflows for AI-Assisted Development](https://developers.redhat.com/articles/2026/04/07/harness-engineering-structured-workflows-ai-assisted-development) — Red Hat's enterprise perspective on harness engineering (April 7, 2026): AI writes better code when you design the environment it works in. Emphasizes structured context over free-form tickets, expanding the agent's toolbox through MCP integrations (CI status, deployment logs, runtime metrics) as real data sources, and a four-pillar model (vibes, specs, skills, agents) for organizing how humans and agents collaborate.\n- [Harness engineering for coding agent users](https://martinfowler.com/articles/harness-engineering.html) — Birgitta Böckeler's systematic mental model (April 2026) for coding-agent harnesses, framing them as feedforward guides plus feedback sensors that self-correct before output reaches human eyes. Distinguishes computational controls (linters, tests) from inferential ones (LLM-as-judge), and argues that harnessability should become a first-class criterion in technology and architecture decisions.\n- [A Practical Guide to Building AI Agents](https://openai.com/business/guides-and-resources/a-practical-guide-to-building-ai-agents/) — OpenAI's April 2026 comprehensive guide distilling production deployment patterns into actionable best practices: single-agent vs. multi-agent orchestration (manager vs. decentralized handoffs), tool design for many-to-many agent-tool relationships, and layered guardrail patterns combining input validation, output filtering, tool-risk ratings, and human-intervention triggers.\n- [An Update on Recent Claude Code Quality Reports](https://www.anthropic.com/engineering/april-23-postmortem) — Anthropic's transparent April 2026 postmortem tracing Claude Code quality degradation to three independent harness-level changes: a default reasoning-effort downgrade, a caching-optimization bug that continuously dropped thinking history from stale sessions, and an overly aggressive verbosity-limiting system prompt. Essential reading for understanding how seemingly minor harness adjustments — prompt wording, cache headers, and default parameters — can compound into visible agent regressions, and for the rigorous diagnostic process required to isolate them.\n- [Agent Harness Design: 3 Patterns for Harnessing Claude's Intelligence](https://claude.com/blog/harnessing-claudes-intelligence) — Anthropic's April 2026 design guide distilling harness engineering into three actionable patterns: build on tools Claude already knows, remove harness assumptions as capabilities improve, and set UX/cost/safety boundaries carefully. A practical complement to the \"agent = model + harness\" framing that helps teams decide what scaffolding to keep, add, or remove over time.\n- [Code as Agent Harness](https://arxiv.org/abs/2605.18747) — May 2026 survey framing code as the basis for agent infrastructure rather than merely output: it unifies harness interface, mechanisms, and multi-agent scaling through shared code artifacts, and surfaces open challenges from verification under incomplete feedback to regression-free improvement.\n- [Harness Engineering: How to Build Reliable AI Agents by Engineering the System, Not the Model](https://www.deepset.ai/blog/harness-engineering) — deepset's May 2026 synthesis of agent reliability as a harness problem: a failure-classification framework (context, constraint, verification, planning failures) that maps each failure mode to the right harness component, and a concrete demonstration that harness-only changes can move agents 20+ ranking positions without swapping the model.\n- [What makes a harness a harness: necessary and sufficient conditions for an agent harness](https://arxiv.org/abs/2606.10106) — June 2026 constitutive definition of an agent harness as a runtime layer with four necessary and sufficient elements: an agent loop, a tool interface, context management, and control mechanisms. Applied to Claude Code, Codex CLI, Aider, Cline, OpenHands, and SWE-agent, it provides a rigorous inclusion test for distinguishing harnesses from generators, guardrails, or plain tool wrappers.\n- [Architectural Design Decisions in AI Agent Harnesses](https://arxiv.org/abs/2604.18071) — April 2026 empirical study of 70 public agent systems across five recurring dimensions (subagent architecture, context management, tool systems, safety mechanisms, orchestration) that synthesizes five architectural patterns. The comparative research package turns harness selection from a framework popularity contest into a reasoned comparison of design trade-offs.\n- [RUCAIBox/awesome-agent-harness](https://github.com/RUCAIBox/awesome-agent-harness) — RUCAIBox's survey paper and curated reading list on *Agent Systems with Harness Engineering*, mapping harness design across agent workflows, memory systems, skill libraries, and multi-agent orchestration with 500+ references. The clearest academic complement to vendor-specific harness engineering posts. ![Stars](https://img.shields.io/github/stars/RUCAIBox/awesome-agent-harness?style=flat-square\u0026label=★\u0026color=yellow)\n- [Tuning the harness, not the model: a Nemotron 3 Ultra playbook](https://blog.langchain.com/tuning-the-harness-not-the-model-a-nemotron-3-ultra-playbook) — LangChain's July 2026 playbook showing how harness-only tuning brought Nemotron 3 Ultra within one point of Opus 4.8 on Deep Agents at roughly one-tenth the cost ($4.48 vs $43.48). The clearest recent demonstration that evals are the training data for harness work and that fit — not raw model capability — determines how much quality reaches the task.\n- [lopopolo/harness-engineering](https://github.com/lopopolo/harness-engineering) — Ryan Lopopolo's anthology, field guide, and agent context bundle for harness engineering: it reframes the harness as the environment that carries an organization's nonfunctional requirements, with reusable `AGENTS.md`/`CLAUDE.md` artifacts, playbooks, evals, and domain modeling docs. The most systematic open-source synthesis of how to make organizational judgment cumulative across agent-maintained repositories. ![Stars](https://img.shields.io/github/stars/lopopolo/harness-engineering?style=flat-square\u0026label=★\u0026color=yellow)\n\n---\n\n## Design Primitives\n\nHarness components organized by the problem they solve, not by vendor.\n\n### Agent Loop\n\n- [ReAct: Synergizing Reasoning and Acting in Language Models](https://arxiv.org/abs/2210.03629) — The foundational paper defining the Thought/Action/Observation loop structure that underlies virtually every agent harness. Required reading for understanding why the loop is structured the way it is and where each harness component maps onto the reasoning-acting cycle.\n- [Unrolling the Codex Agent Loop](https://openai.com/index/unrolling-the-codex-agent-loop/) — The canonical decomposition of what happens inside one agent loop iteration: observe, plan, act, verify.\n- [LangGraph — Low Level Concepts](https://langchain-ai.github.io/langgraph/concepts/low_level/) — Models the agent loop explicitly as a directed graph with typed state, conditional edges, and checkpointing. The most concrete engineering treatment of loop control flow: how to implement termination conditions, branch on tool results, and persist mid-loop state for resumption.\n- [Unlocking the Codex Harness: How We Built the App Server](https://openai.com/index/unlocking-the-codex-harness/) — OpenAI's engineering deep-dive into the Item/Turn/Thread protocol (JSON-RPC/JSONL over stdio) that exposes the Codex harness to every client surface. The most direct first-party account of why approval flows, streaming diffs, and thread persistence demand a purpose-built protocol — and why MCP's tool-oriented model proved insufficient for these requirements.\n- [Hooks – Codex](https://developers.openai.com/codex/hooks) — OpenAI's lifecycle-hook framework for Codex: inject deterministic scripts at `SessionStart`, `PreToolUse`, `PostToolUse`, and other loop events to enforce guardrails, audit actions, and customize agent behavior without relying on prompt-level trust. A concrete reference for programmable harness governance.\n- [Extended Thinking — Claude API Docs](https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking) — The harness-critical reference for integrating extended thinking into agent loops: `budget_tokens` controls reasoning depth per turn, thinking blocks **must be preserved** when passing tool results back (omitting them silently breaks multi-step reasoning), and thinking mode cannot change mid-turn. Essential before wiring extended thinking into any tool-use loop.\n- [Getting started with loops](https://claude.com/blog/getting-started-with-loops) — Anthropic's June 2026 practical taxonomy of agent loops: turn-based, goal-based (`/goal`), time-based (`/loop`, `/schedule`), and proactive loops. The framework for matching loop primitive to task shape — and the emphasis on deterministic stop conditions and token budgets — makes it a concise reference for choosing the right loop abstraction instead of defaulting to a single conversational turn cycle.\n- [Improving Deep Agents with Harness Engineering](https://blog.langchain.com/improving-deep-agents-with-harness-engineering/) — LangChain's case study showing harness-only changes moved their coding agent from rank 30 to top 5 on Terminal Bench 2.0 with no model swap: structured verification loops, context injection (directory maps + time budget warnings), loop-detection middleware, and a \"reasoning sandwich\" concentrating maximum thinking at planning and verification phases. The most concrete published demonstration that harness design is the primary performance lever, not model capability.\n- [Loop Engineering](https://github.com/cobusgreyling/loop-engineering) — Practical design system for agent loops with seven production patterns, cross-tool starter kits, and CLI tools that score readiness, scaffold state, estimate cost, detect drift, and isolate worktrees. The clearest open-source resource for moving from one-off prompting to durable, observable agent loops. ![Stars](https://img.shields.io/github/stars/cobusgreyling/loop-engineering?style=flat-square\u0026label=★\u0026color=yellow)\n- [Life-Harness](https://github.com/Tianshi-Xu/Life-Harness) — Official implementation of a lifecycle-aware runtime harness that improves frozen LLM agents by adapting the model-environment interface across four layers: environment contract, procedural skills, action realization, and trajectory regulation. The key result is that harness-side adaptation transfers across 18 model backbones, proving that many agent failures are interface mismatches rather than reasoning deficits. ![Stars](https://img.shields.io/github/stars/Tianshi-Xu/Life-Harness?style=flat-square\u0026label=★\u0026color=yellow)\n- [How Middleware Lets You Customize Your Agent Harness](https://blog.langchain.com/how-middleware-lets-you-customize-your-agent-harness/) — Introduces AgentMiddleware: six composable hooks (`before_agent`, `before_model`, `wrap_model_call`, `wrap_tool_call`, `after_model`, `after_agent`) that intercept every stage of the agent loop. Enables deterministic policy enforcement (PII redaction that can't be trusted to prompts), dynamic tool injection, mid-task model swapping, and production patterns (retry, fallback, HITL interrupts) without modifying core agent logic — the reference design for cross-cutting harness concerns that shouldn't be baked into individual agents.\n- [Agents Learn Their Runtime: Interpreter Persistence as Training-Time Semantics](https://arxiv.org/abs/2603.01209) — Controlled experiment isolating interpreter state persistence as an independent training variable. The harness finding: mismatching your runtime persistence mode to the model's training-time semantics produces either 80% missing-variable errors (model expects state that doesn't persist) or 3.5× token overhead (model redundantly recomputes state it expects to already have). Persistence is a learned semantic that must be honored at deployment, not a free runtime choice.\n- [Real-Time Deadlines Reveal Temporal Awareness Failures in LLM Strategic Reasoning](https://arxiv.org/abs/2601.13206) — Demonstrates that temporal awareness (handling deadlines and time constraints) appears orthogonal to reasoning capability: explicit temporal feedback in the agent loop significantly improves LLM performance on deadline-constrained tasks. Indicates temporal semantics as a learned behavior that must be integrated into harness-level context (current time, deadlines, time budgets) rather than assumed from capability alone.\n- [A Scheduler-Theoretic Framework for LLM Agent Execution](https://arxiv.org/abs/2604.11378) — April 2026 systematic analysis of 70 open-source LLM agent projects showing 60% adopt the Agent Loop pattern. Proposes a formal scheduler framework that maps execution patterns (Agent Loop, Event-driven, State-machine, Graph/flow, Hybrid) onto a unified control model, making the controllability/expressiveness/implementability trade-offs explicit. Essential reading for choosing the right loop architecture rather than defaulting to the simplest pattern.\n- [Confucius Code Agent (CCA)](https://github.com/facebookresearch/cca-swebench) — February 2026 production-grade coding agent from Meta/Harvard built on the Confucius SDK, which structures harness design around three perspectives: Agent Experience (AX), User Experience (UX), and Developer Experience (DX). Features a unified orchestrator with advanced context management, persistent note-taking for cross-session learning, and a meta-agent that automates build-test-improve cycles. Achieves 59% Resolve@1 on SWE-Bench-Pro, exceeding prior research and commercial baselines.\n- [The Design Space of Today's and Future AI Agent Systems](https://arxiv.org/abs/2604.14228) — April 2026 reverse-engineering of Claude Code's architecture revealing five-stage progressive compaction (budget reduction → snip → microcompact → context collapse → auto-compact), subagent isolation with rebuilt permission contexts, and a 27-event-type hook pipeline. The most detailed public analysis of a production agent loop's internal design decisions — essential for understanding how context pressure, safety, and delegation are handled at scale.\n- [deepclaude](https://github.com/aattaran/deepclaude) — Ports Claude Code's full agent loop to DeepSeek V4 Pro and other Anthropic-compatible backends while preserving the same UX. The strongest practical evidence that loop architecture — not model identity — determines agent behavior, and a concrete starting point for building backend-agnostic harnesses. ![Stars](https://img.shields.io/github/stars/aattaran/deepclaude?style=flat-square\u0026label=★\u0026color=yellow)\n- [The Coding Harness Behind GitHub Copilot in VS Code](https://code.visualstudio.com/blogs/2026/05/15/agent-harnesses-github-copilot-vscode) — VS Code team's breakdown of the coding harness behind GitHub Copilot: three core loop responsibilities (context assembly, tool exposure, tool execution), multi-provider model routing across Anthropic, Google, OpenAI, xAI, and Mistral, and the VSC-Bench eval suite with PR-gated assessment. The clearest published account of how a major product treats harness changes as first-class code review criteria — \"the model is the engine, the harness is the car.\"\n- [statewright](https://github.com/statewright/statewright) — State machine guardrails that constrain which tools an agent can call in each phase of a workflow, turning open-ended loops into deterministic state transitions. The research result is striking: local models went from 2/10 to 10/10 passing on a SWE-bench subset purely by shrinking the tool space, proving that loop structure — not model size — is the binding constraint. ![Stars](https://img.shields.io/github/stars/statewright/statewright?style=flat-square\u0026label=★\u0026color=yellow)\n- [Introducing dynamic workflows in Claude Code](https://claude.com/blog/introducing-dynamic-workflows-in-claude-code) — Anthropic's May 2026 introduction to dynamic parallel subagent orchestration: Claude generates JavaScript orchestration scripts that fan out work to tens or hundreds of parallel subagents with adversarial verification, converging on answers for tasks like the 750k-line Bun Zig-to-Rust port. The key harness insight is that the plan lives in executable code rather than the model's context window, scaling the agent loop to work that would otherwise exceed a single context window.\n- [AgentSPEX](https://github.com/ScaleML/AgentSPEX) — UIUC's open-source specification and execution language for LLM-agent workflows: declarative YAML with typed steps, branching, loops, and explicit state management, backed by a Docker sandbox with 50+ MCP tools, checkpointing, and trajectory logging. A concrete reference for turning ad-hoc agent loops into version-controlled, reproducible harness artifacts. ![Stars](https://img.shields.io/github/stars/ScaleML/AgentSPEX?style=flat-square\u0026label=★\u0026color=yellow)\n\n### Planning \u0026 Task Decomposition\n\n- [Run Long-Horizon Tasks with Codex](https://developers.openai.com/blog/run-long-horizon-tasks-with-codex/) — Introduces milestone-based planning artifacts (Plan.md, Implement.md) as harness-level state.\n- [Harness Design for Long-Running Application Development](https://www.anthropic.com/engineering/harness-design-long-running-apps) — Multi-session planning, progress tracking, and the role of persistent planning documents.\n- [Plan-and-Execute Agents](https://blog.langchain.com/plan-and-execute-agents/) — The canonical engineering write-up separating planning from execution as distinct harness layers: a planner LLM generates the step list once; an executor agent works through it, replanning only when needed. Defines the pattern that most modern task-decomposition harnesses follow.\n- [microsoft/TaskWeaver](https://github.com/microsoft/TaskWeaver) — Code-first task decomposition framework with a planner/executor split and a plugin system for injecting domain knowledge into the planning layer. The most complete reference implementation of plan-then-execute with stateful task tracking. ![Stars](https://img.shields.io/github/stars/microsoft/TaskWeaver?style=flat-square\u0026label=★\u0026color=yellow)\n- [LATS: Language Agent Tree Search](https://arxiv.org/abs/2310.04406) — Unifies reasoning, acting, and planning via Monte Carlo Tree Search over agent trajectories. Directly informs harness design: external tool feedback as tree-search signals, trajectory backtracking on failure, and depth-bounded exploration make this the most actionable planning research for harnesses with real environment interaction.\n- [Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering](https://arxiv.org/abs/2602.01465) — Demonstrates specialized harness patterns for coordinating heterogeneous agent teams (planner, coder, reviewer, executor) on software engineering tasks. Shows how role-specific agents with different model sizes and tool access produce better outcomes than single-agent approaches, with concrete metrics on task decomposition effectiveness.\n- [Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks](https://arxiv.org/abs/2503.09572) — Modular framework separating high-level planning from low-level execution through synthetic data generation and explicit structured planning. Achieves 57.58% success on WebArena-Lite and 81.36% on WebVoyager. The key harness insight is that planner and executor can be specialized independently — different model sizes, tool access, and reasoning budgets for each layer — improving overall reliability on tasks exceeding context window limits.\n- [Choosing the Right Multi-Agent Architecture](https://blog.langchain.com/choosing-the-right-multi-agent-architecture/) — Decision framework for four multi-agent patterns (subagents, skills, handoffs, router) with concrete performance data: subagents process 67% fewer tokens than skills in multi-domain scenarios because context isolation prevents cross-domain bloat. The five-dimension matching table (distributed development, parallelization, multi-hop, user interaction, latency) is the most actionable published guide for deciding when a topology change — not a model change — is the right lever for a performance problem.\n- [Multi-Agent Workflows Often Fail. Here's How to Engineer Ones That Don't.](https://github.blog/ai-and-ml/generative-ai/multi-agent-workflows-often-fail-heres-how-to-engineer-ones-that-dont/) — GitHub's February 24, 2026 distillation of a failure pattern most harnesses eventually rediscover: multi-agent systems behave like distributed systems, so every handoff needs typed schemas, constrained action schemas, and explicit boundary validation. Worth including because it turns \"add more agents\" from a vibe into an interface design problem you can actually reason about.\n- [Effective Harnesses for Long-Running Agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) — Anthropic's pattern for maintaining agent progress across multiple context windows: an initializer agent sets up the environment once and hands off to a coding agent that makes incremental progress each session. The structured handoff mechanism — feature lists, git commits, and test gates as cross-session state — is the reference design for any harness where a task exceeds a single context window and naïve restarts lose accumulated progress.\n- [Task-Adaptive Multi-Agent Orchestration (AdaptOrch)](https://arxiv.org/abs/2602.16873) — February 2026 framework that dynamically selects orchestration topology (parallel, sequential, hierarchical, or hybrid) based on task dependency graphs rather than fixed pipeline architecture. Demonstrates that topology choice is a harness-level lever that can improve performance 12–23% over model selection alone.\n- [Task-Decoupled Planning for Long-Horizon Agents (TDP)](https://arxiv.org/abs/2601.07577) — January 2026 planning framework that combines task decomposition with modular agent design: a Supervisor decomposes tasks into a dependency graph, Planner \u0026 Executor agents solve each decoupled sub-task node independently, and a Self-Revision module updates the graph after execution. The key harness insight is that decoupling planning from execution at the sub-task level enables localized replanning without cascading failures across the entire task chain.\n\n### Context Delivery \u0026 Compaction\n\n- [Harness Engineering](https://openai.com/index/harness-engineering/) — How to structure context windows for agents: what to include, what to exclude, and how context shape affects agent behavior.\n- [Effective Context Engineering for AI Agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) — Anthropic's systematic guide to managing the full context state—system prompts, tools, MCP, and message history—as a finite, curated resource. Reframes harness design as \"what configuration of context produces the desired behavior?\" rather than just prompt wording.\n- [Compaction — Claude API Docs](https://platform.claude.com/docs/en/build-with-claude/compaction) — Anthropic's reference for server-side context compaction: automatically summarizes older context when approaching the window limit. Reduced token consumption by 84% in a 100-turn web search eval while allowing agents to complete workflows that would otherwise hit context limits.\n- [LLMLingua](https://github.com/microsoft/LLMLingua) — Microsoft Research's prompt compression toolkit (up to 20x compression, minimal performance loss) that can be embedded as a preprocessing step in the context delivery layer. LLMLingua-2 adds 3–6x speed gains, making it viable for latency-sensitive agent loops. ![Stars](https://img.shields.io/github/stars/microsoft/LLMLingua?style=flat-square\u0026label=★\u0026color=yellow)\n- [Prompt Caching — Claude API Docs](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) — The most effective harness-level cost lever: cache repeated system prompts, tool definitions, and long documents across requests. Explains where to place `cache_control` breakpoints for maximum reuse across multi-turn agent sessions.\n- [Autonomous Context Compression](https://blog.langchain.com/autonomous-context-compression/) — Shifts context compression from harness-controlled (compacting at a fixed token threshold) to agent-controlled: agents call a dedicated tool to trigger compression when strategically appropriate — between tasks or before consuming large inputs. Eliminates the failure mode where reactive-at-limit compaction interrupts agents mid-subtask and corrupts in-flight reasoning state.\n- [Active Context Compression: Autonomous Memory Management in LLM Agents](https://arxiv.org/abs/2601.07190) — Proposes a \"Focus Agent\" architecture where the agent autonomously decides when to consolidate interaction history into a persistent Knowledge block and prune raw context — shifting compression from a harness-enforced policy to a model-controlled action. Produces 22.7% token reduction with no accuracy loss on long-horizon tasks; the core contribution is making the compression unit semantically coherent (the agent decides what knowledge is worth preserving) rather than mechanically token-budget-driven.\n- [context-mode](https://github.com/mksglu/context-mode) — MCP server that intercepts raw tool output before it enters the context window, sandboxing bulky data (Playwright snapshots, GitHub issues, logs) outside the LLM and retrieving only relevant fragments via BM25 when needed. The \"think in code\" paradigm — replacing ten file-read tool calls with one script execution — is a concrete harness pattern for turning context pressure into a programming problem rather than a compression problem. ![Stars](https://img.shields.io/github/stars/mksglu/context-mode?style=flat-square\u0026label=★\u0026color=yellow)\n- [Making Agent-Friendly Pages with Content Negotiation](https://vercel.com/blog/making-agent-friendly-pages-with-content-negotiation) — Vercel's February 3, 2026 implementation guide for serving `text/markdown` when agents request it via `Accept: text/markdown`, while preserving the same human-facing HTML URL. This is a real harness primitive, not just a docs trick: it removes boilerplate before it ever enters the context window and gives agents cleaner, cheaper inputs without custom scrapers.\n- [A-RAG: Scaling Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces](https://arxiv.org/abs/2602.03442) — Reframes RAG as a harness tool-design problem: instead of injecting retrieved documents into context at pipeline time, expose three retrieval tools (keyword search, semantic search, chunk read) and let the agent pull information incrementally as each reasoning step requires it. The key harness decision is architectural — retrieval becomes a tool call in the agent loop, not a preprocessing step — which means the agent's reasoning can adaptively narrow scope rather than processing everything injected upfront.\n- [LLM Readiness Harness: Evaluation, Observability, and CI Gates for LLM/RAG Applications](https://arxiv.org/abs/2603.27355) — Structured framework for building production-grade evaluation harnesses: evaluation gates that block deployment, observability instrumentation that tracks all agent decisions, and CI integration patterns that catch regressions before they reach users. Essential reading for organizations deploying multiple agents in parallel where a single harness failure can cascade.\n- [ByteRover: Agent-Native Memory Through LLM-Curated Hierarchical Context](https://arxiv.org/abs/2604.01599) — LLM-curated hierarchical context management for agents where the model itself learns to weight information importance across multiple hierarchy levels. Reduces token overhead through learned relevance filtering without sacrificing comprehension. Directly applicable to any harness where context budget is the limiting factor — letting the model curate what belongs in active memory vs. what can be retrieved on-demand.\n- [Claude Code Compaction: How Context Compression Works](https://okhlopkov.com/claude-code-compaction-explained/) — March 2026 deep-dive into Claude Code's automatic compaction mechanism: what survives (current task, recent errors, file names) vs. what gets lost (initial instructions, intermediate decisions, style rules). Key harness insight: never rely on compaction for critical rules — move them to `CLAUDE.md` where they live in the system prompt and survive any compression. Essential practical guidance for anyone running long-session agents.\n- [Token Savior](https://github.com/Mibayy/token-savior) — MCP server that indexes codebases by symbol (functions, classes, call graphs) so agents navigate by pointer instead of reading whole files, cutting active tokens by 77% and benchmark wall time by 76%. Demonstrates that context delivery for coding agents is a navigation problem, not just a compression problem. ![Stars](https://img.shields.io/github/stars/Mibayy/token-savior?style=flat-square\u0026label=★\u0026color=yellow)\n- [Trellis](https://github.com/mindfold-ai/Trellis) — Replaces the bloated `CLAUDE.md` pattern with a progressive spec system: agents load only the standards, task PRDs, and session journals relevant to the current step. The cross-platform adapter layer turns vendor-specific harness configuration into a portable team practice rather than a per-tool hack. ![Stars](https://img.shields.io/github/stars/mindfold-ai/Trellis?style=flat-square\u0026label=★\u0026color=yellow)\n- [OpenViking](https://github.com/volcengine/OpenViking) — ByteDance's context database for AI agents that unifies memory, resources, and skills through a filesystem paradigm, enabling hierarchical context delivery where agents pull only the paths they need instead of receiving bloated monolithic prompts. The self-evolving layer that restructures context based on usage patterns makes it a rare example of context infrastructure that improves autonomously rather than requiring constant manual curation. ![Stars](https://img.shields.io/github/stars/volcengine/OpenViking?style=flat-square\u0026label=★\u0026color=yellow)\n- [DESIGN.md](https://github.com/google-labs-code/design.md) — Google Labs' specification for describing visual identity systems to coding agents: machine-readable design tokens (YAML front matter) combined with human-readable design rationale (markdown prose) give agents a persistent, structured understanding of design constraints without requiring custom tool chains. ![Stars](https://img.shields.io/github/stars/google-labs-code/design.md?style=flat-square\u0026label=★\u0026color=yellow)\n- [codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp) — High-performance code intelligence MCP server that full-indexes repositories into a persistent knowledge graph via tree-sitter AST analysis across 66 languages. Replaces dozens of file-read/grep cycles with sub-millisecond structured queries, cutting active tokens by 120× and turning codebase navigation from a context-pressure problem into a pointer-chasing problem. ![Stars](https://img.shields.io/github/stars/DeusData/codebase-memory-mcp?style=flat-square\u0026label=★\u0026color=yellow)\n- [Mirage](https://github.com/strukto-ai/mirage) — Mounts S3, Slack, Gmail, GitHub, and Redis side-by-side as a single virtual filesystem so agents interact with every backend through familiar bash commands instead of learning N distinct APIs. The key harness insight: LLMs are already fluent in `grep`, `cat`, and `cp` — leveraging that vocabulary eliminates tool-schema bloat and makes cross-service pipelines compose as naturally as local shell scripts. ![Stars](https://img.shields.io/github/stars/strukto-ai/mirage?style=flat-square\u0026label=★\u0026color=yellow)\n- [dirac](https://github.com/dirac-run/dirac) — Coding agent harness optimized for surgical context curation and API cost reduction: Hash Anchored edits, massively parallel operations, and AST manipulation combine to cut costs 50–80% while improving code quality. Demonstrates that precise context delivery — not just bulk compression — is the right lever for efficient coding agents. ![Stars](https://img.shields.io/github/stars/dirac-run/dirac?style=flat-square\u0026label=★\u0026color=yellow)\n- [MinishLab/semble](https://github.com/MinishLab/semble) — Code search primitive that replaces grep+read cycles with natural-language retrieval, cutting active tokens by ~98% while keeping 99% of a transformer-based retriever's accuracy. Ships as an MCP server and CLI, runs on CPU with zero external dependencies — the right drop-in for any coding agent harness struggling with context pressure. ![Stars](https://img.shields.io/github/stars/MinishLab/semble?style=flat-square\u0026label=★\u0026color=yellow)\n- [harness-experimental](https://github.com/hoangnb24/harness-experimental) — Repository-level operating harness that turns any software repo into an agent-ready workspace: structured `AGENTS.md`, `HARNESS.md`, and `FEATURE_INTAKE.md` give agents the missing project context — where to start, what the product contract says, how risky the change is, and which decisions future agents should inherit. The most concrete open-source implementation of \"coding agents need better repositories, not just better prompts.\" ![Stars](https://img.shields.io/github/stars/hoangnb24/harness-experimental?style=flat-square\u0026label=★\u0026color=yellow)\n- [headroom](https://github.com/chopratejas/headroom) — Compresses tool outputs, logs, files, and RAG chunks before they enter the context window, cutting active tokens by 60–95% without changing answers. Ships as a library, proxy, and MCP server — the right drop-in layer for any harness where bulky tool returns are the primary context pressure source. ![Stars](https://img.shields.io/github/stars/chopratejas/headroom?style=flat-square\u0026label=★\u0026color=yellow)\n- [Context7](https://github.com/upstash/context7) — MCP server and CLI that injects up-to-date, version-specific library documentation directly into agent context, eliminating hallucinated APIs and outdated code examples caused by stale training data. Ships as both a `ctx7` command-line tool and an MCP server with `resolve-library-id` and `query-docs` tools. ![Stars](https://img.shields.io/github/stars/upstash/context7?style=flat-square\u0026label=★\u0026color=yellow)\n- [Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning](https://arxiv.org/abs/2605.15315) — Decomposes code relevance into two interpretable dimensions — semantic evidence and dependency support — rather than collapsing all retention decisions into a single score. Saves up to 31% more tokens on multi-turn coding agent tasks while improving Exact Match by up to +3.5, demonstrating that coding-agent context pruning needs domain-specific rubrics rather than generic compression.\n- [OpenWiki](https://github.com/langchain-ai/openwiki) — LangChain's CLI that writes and maintains agent-readable wikis for codebases or purpose memory, turning documentation drift into a versioned, automatable harness artifact. Emits Google Open Knowledge Format bundles so curated context stays portable across agents and can be kept fresh via CI. ![Stars](https://img.shields.io/github/stars/langchain-ai/openwiki?style=flat-square\u0026label=★\u0026color=yellow)\n- [PRO-LONG](https://github.com/alexisfox7/PRO-LONG) — Programmatic memory framework for long-horizon agents: the harness appends all observations to a structured log and lets the agent search it with code instead of relying on fixed summarization or compaction policies. Achieves 4.2–5.8× token reduction and matches or exceeds specialized harnesses on ARC-AGI-3, showing that long-horizon context management can be a programmable retrieval problem rather than a compression problem. ![Stars](https://img.shields.io/github/stars/alexisfox7/PRO-LONG?style=flat-square\u0026label=★\u0026color=yellow)\n- [Graft](https://github.com/NanoNets/Graft) — Builds a local, regenerable graph of plain-English system explanations and code relationships, then rides along inside Claude Code, Cursor, Codex, and Gemini via MCP and statusline hooks so the agent stops rediscovering the repo every session. The published SWE-bench Verified and efficiency benchmarks make it the clearest recent demonstration that context delivery for coding agents is a navigation problem, not just a compression problem. ![Stars](https://img.shields.io/github/stars/NanoNets/Graft?style=flat-square\u0026label=★\u0026color=yellow)\n- [ktx](https://github.com/Kaelio/ktx) — Self-improving executable context layer for data and analytics agents: it ingests warehouses, BI tools, and wikis to build a semantic layer with approved metrics, joinable columns, and resolved fan/chasm traps, then serves the result to Claude Code, Codex, and Cursor through MCP. Fills the gap where general-purpose agents invent metric logic on every question — it turns warehouse querying from a prompt-guessing problem into a governed, retrievable context problem. ![Stars](https://img.shields.io/github/stars/Kaelio/ktx?style=flat-square\u0026label=★\u0026color=yellow)\n\n### Tool Design\n\n- [Writing Effective Tools for Agents](https://www.anthropic.com/engineering/writing-effective-tools-for-agents) — Tool naming, schema design, error messages, and return value conventions that make agents more reliable.\n- [Tool Use — Claude API Docs](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview) — Authoritative reference for client vs. server tool execution models, strict schema enforcement, and tool_result error signaling. The distinction between client-side and server-side tool execution is a foundational harness architecture decision.\n- [Function Calling — OpenAI Docs](https://platform.openai.com/docs/guides/function-calling) — Defines the de facto industry-standard JSON Schema conventions for tool definitions and parallel function calling. Essential reading before designing a tool interface that needs to work across multiple models.\n- [Tool Annotations as Risk Vocabulary](https://blog.modelcontextprotocol.io/posts/2026-03-16-tool-annotations/) — The MCP team's definitive post on the four tool annotation hints (`readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`) as inputs to harness permission decisions, not enforced contracts. The \"lethal trifecta\" — private data access + untrusted content exposure + external communication — is the most actionable framing for why single-tool safety analysis misses the risk that emerges from tool *combinations*.\n- [outlines](https://github.com/dottxt-ai/outlines) — Constrains token sampling via regex/CFG/JSON Schema at the decoding layer, guaranteeing structured output without model fine-tuning. The right solution when you need OpenAI Structured Outputs-equivalent reliability from a locally deployed or open-weight model. ![Stars](https://img.shields.io/github/stars/dottxt-ai/outlines?style=flat-square\u0026label=★\u0026color=yellow)\n- [instructor](https://python.useinstructor.com/) — Maps Pydantic models directly to structured LLM extraction with built-in retry and validation-error feedback loops. Turns tool call output parsing from ad-hoc JSON handling into type-safe data models, eliminating an entire class of harness parsing bugs. ![Stars](https://img.shields.io/github/stars/instructor-ai/instructor?style=flat-square\u0026label=★\u0026color=yellow)\n- [SkillTester: Benchmarking Utility and Security of Agent Skills](https://arxiv.org/abs/2603.28815) — Framework for evaluating agent skills on three dimensions (capability, robustness, security) before deployment. Directly addresses the harness problem of skill sprawl: as agents gain access to more tools, the combinatorial explosion of failure modes becomes unmanageable without systematic verification. The 86-task benchmark across 11 domains provides reference metrics for skill quality.\n- [AutoHarness: Improving LLM Agents by Automatically Synthesizing a Code Harness](https://arxiv.org/abs/2603.03329) — Google DeepMind technique that uses code synthesis to auto-generate runtime constraint harnesses from tool schemas and task specifications. Gemini-2.5-Flash + AutoHarness outperforms Gemini-2.5-Pro and GPT-5.2-High on TextArena games by eliminating illegal moves through learned harness policies. Shifts constraint enforcement from static (schema validation) to dynamic (synthesized code guards) — a reference pattern for learning-based behavioral guardrails.\n- [Scaling Parallel Tool Calling for Efficient Deep Research](https://arxiv.org/abs/2602.07359) — February 2026 analysis of how parallel tool calling reduces latency in multi-step agent workflows. Demonstrates that concurrent tool execution (rather than sequential observe→act loops) is the key efficiency lever for deep-research harnesses where each step may invoke search, browse, and compute tools simultaneously. Essential for designing low-latency agent loops without sacrificing reasoning depth.\n- [EigentSearch-Q+](https://arxiv.org/abs/2604.07927) — April 2026 framework for deep-research agents using dedicated reasoning tools (plan_next_searches, select_query_and_search, extract_relevant_details, analyze_search_progress) that externalize intermediate decisions as typed tool arguments. Inspired by Anthropic's think-tool paradigm, Q+ makes cognitive scaffolding explicit and auditable — bridging classic information-retrieval strategies with structured model-driven tool invocations.\n- [TopoCurate: Modeling Interaction Topology for Tool-Use Agent Training](https://arxiv.org/abs/2603.01714) — March 2026 framework that models interaction topology — the structural patterns of how agents invoke, chain, and conditionally branch between tools — as a first-class training signal. Rather than treating tool use as isolated function calls, TopoCurate learns topological priors from expert trajectories, improving generalization to novel tool combinations and multi-step orchestration patterns. Directly applicable to harnesses where tool topology (not just tool availability) determines task success.\n- [Design Patterns for Deploying AI Agents with Model Context Protocol](https://arxiv.org/abs/2603.13417) — March 2026 field report from an enterprise MCP deployment identifying three protocol-level gaps that break production: missing identity propagation (who is the request for?), absent adaptive tool budgeting, and unstructured error semantics. The concrete mitigation patterns — JWT-enriched tool calls, per-tool timeout contracts, and standardized error-action mappings — are essential before betting on MCP as your primary tool-integration layer.\n- [tui-use](https://github.com/onesuper/tui-use) — Expands the agent tool surface beyond non-interactive commands: programmable TUI interaction for REPLs, debuggers, and ncurses apps that standard bash can't reach. A concrete harness primitive for any agent that needs to operate interactive CLI tools without building a custom wrapper per program. ![Stars](https://img.shields.io/github/stars/onesuper/tui-use?style=flat-square\u0026label=★\u0026color=yellow)\n- [CLI-Anything](https://github.com/HKUDS/CLI-Anything) — Generates agent-native CLI harnesses for any software, giving agents structured JSON access to applications that were never designed for automation. The CLI-Hub registry and auto-generated SKILL.md files turn tool expansion into a package-manager experience — solving the \"long tail\" of agent tool coverage without a custom wrapper per program. ![Stars](https://img.shields.io/github/stars/HKUDS/CLI-Anything?style=flat-square\u0026label=★\u0026color=yellow)\n- [zerolang](https://github.com/vercel-labs/zerolang) — Experimental graph-first programming language where agents inspect and edit code through a compiler-derived ProgramGraph (node IDs, graph hashes, types, effects, ownership) instead of fragile text patches. Collapses the typical agent loop of edit-format-reparse-check-fix into a single compiler-validated semantic operation — a reference design for making code manipulation a structured tool interface rather than a guesswork-driven text transformation. ![Stars](https://img.shields.io/github/stars/vercel-labs/zerolang?style=flat-square\u0026label=★\u0026color=yellow)\n\n### Skills \u0026 MCP\n\n- [Model Context Protocol](https://modelcontextprotocol.io/introduction) — Anthropic's open protocol for connecting agents to external tools, data sources, and services in a standardized way.\n- [modelcontextprotocol/servers](https://github.com/modelcontextprotocol/servers) — Anthropic's official reference MCP server implementations (GitHub, Slack, Postgres, Puppeteer, etc.). The authoritative source for understanding correct MCP server structure before building your own. ![Stars](https://img.shields.io/github/stars/modelcontextprotocol/servers?style=flat-square\u0026label=★\u0026color=yellow)\n- [microsoft/playwright-mcp](https://github.com/microsoft/playwright-mcp) — Browser automation via accessibility tree snapshots rather than screenshots, dramatically reducing token cost. The canonical example of structured tool output design in an MCP server. ![Stars](https://img.shields.io/github/stars/microsoft/playwright-mcp?style=flat-square\u0026label=★\u0026color=yellow)\n- [Chrome DevTools MCP](https://github.com/ChromeDevTools/chrome-devtools-mcp) — Official Google MCP server that exposes live Chrome debugging surfaces — network analysis, performance profiling, console messages, memory snapshots, and Lighthouse audits — as structured agent tools. The clearest reference for turning browser inspection into a first-class tool interface rather than relying solely on screenshot-driven automation. ![Stars](https://img.shields.io/github/stars/ChromeDevTools/chrome-devtools-mcp?style=flat-square\u0026label=★\u0026color=yellow)\n- [agent-device](https://github.com/callstackincubator/agent-device) — MCP-native control layer for iOS and Android devices: snapshots, semantic targeting, typed client access, diagnostics, and replayable workflows. Fills a critical gap in the mobile-agent harness stack — most tool design assumes desktop or browser surfaces, but real-world agents increasingly need to interact with native mobile apps. ![Stars](https://img.shields.io/github/stars/callstackincubator/agent-device?style=flat-square\u0026label=★\u0026color=yellow)\n- [A2A Protocol](https://github.com/a2aproject/A2A) — Google's open Agent-to-Agent protocol: JSON-RPC over HTTP(S)/SSE with Agent Card service discovery and a task/message/artifact communication model. The emerging standard for cross-framework agent interoperability in multi-agent harnesses. ![Stars](https://img.shields.io/github/stars/a2aproject/A2A?style=flat-square\u0026label=★\u0026color=yellow)\n- [Announcing the Agentic Resource Discovery specification](https://developers.googleblog.com/announcing-the-agentic-resource-discovery-specification/) — Google's June 2026 open specification for publishing, discovering, and verifying AI capabilities across the web via domain-owned catalogs and searchable registries. Adds the missing discovery layer that lets agents find MCP servers, A2A agents, and OpenAPI tools at runtime rather than relying on hardcoded integrations, with trust manifests and namespaced URNs for governance.\n- [MCP Inspector](https://github.com/modelcontextprotocol/inspector) — Interactive debugging UI for MCP servers: inspect tool definitions, send test calls, and validate responses without wiring up a full agent. The essential development tool for anyone building or integrating MCP servers into a harness. ![Stars](https://img.shields.io/github/stars/modelcontextprotocol/inspector?style=flat-square\u0026label=★\u0026color=yellow)\n- [Shell + Skills + Compaction: Tips for Long-Running Agents](https://developers.openai.com/blog/skills-shell-tips) — OpenAI's engineering guide to three production harness primitives: versioned Skill bundles (SKILL.md manifest; routing accuracy improved 73%→85% by adding negative examples), a managed shell container for durable tool execution, and server-side compaction via explicit `/responses/compact` endpoint. The most concrete first-party documentation of skills-based routing and compaction published in 2026.\n- [Composio](https://github.com/ComposioHQ/composio) — Wraps 250+ SaaS APIs (GitHub, Slack, Linear, Notion, etc.) as agent-ready actions with managed OAuth, so tool integration becomes a one-line import rather than a custom harness component per service. The fastest path from \"the agent needs to call an external API\" to a production-grade, authenticated tool. ![Stars](https://img.shields.io/github/stars/ComposioHQ/composio?style=flat-square\u0026label=★\u0026color=yellow)\n- [oomol-lab/open-connector](https://github.com/oomol-lab/open-connector) — Open-source auth gateway connecting 1000+ SaaS providers to AI agents through SDK, CLI, MCP, HTTP, and OpenAPI. Treats external-tool onboarding as a unified, self-hosted harness layer: one credential and access model covers direct SDK calls, MCP servers, and OpenAPI discovery, so teams don't rebuild auth per integration. ![Stars](https://img.shields.io/github/stars/oomol-lab/open-connector?style=flat-square\u0026label=★\u0026color=yellow)\n- [MCP Streamable HTTP Transport](https://modelcontextprotocol.io/specification/2025-11-25/basic/transports) — The transport that replaced HTTP+SSE in the 2025-11-25 spec, enabling MCP servers to run as remote services rather than local processes. Servers handle multiple client connections using HTTP POST (for client→server messages) and optional GET (for server→client SSE streams). The key harness architecture decision: Streamable HTTP unlocks remote MCP deployment but introduces session management complexity — stateful `Mcp-Session-Id` headers fight with load balancers and horizontal scaling, which the 2026 roadmap aims to resolve by decoupling sessions from the transport layer.\n- [The 2026 MCP Roadmap](https://blog.modelcontextprotocol.io/posts/2026-mcp-roadmap/) — The MCP team's roadmap for the next spec cycle: horizontal-scaling transport without stateful session constraints, `.well-known` discovery for capability advertisement without live connections, Tasks primitive with retry/expiry semantics, and enterprise extensions (audit trails, SSO, gateway behavior). Essential reading before investing heavily in MCP server infrastructure — the transport and discovery changes in particular affect how clients locate and connect to servers.\n- [The 2026-07-28 MCP Specification Release Candidate](https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/) — The largest revision of MCP since launch: a stateless protocol core drops the `initialize` handshake and `Mcp-Session-Id`, the new `ext-*` extension framework formalizes Tasks and MCP Apps, and a twelve-month deprecation policy gives harness builders a stable target. Essential reading before designing remote MCP server infrastructure that must survive load balancers and horizontal scaling.\n- [Developer's Guide to AI Agent Protocols](https://developers.googleblog.com/en/developers-guide-to-ai-agent-protocols/) — Google's survey of six standardized agent interoperability protocols, each solving a distinct harness integration problem: MCP (tool/data connectivity), A2A (inter-agent routing via Agent Card discovery at well-known URLs), UCP (commerce workflows), AP2 (payment authorization with spend limits), A2UI (agent-driven dynamic UIs), AG-UI (streaming event format). The most practical map of which protocol to choose when an agent needs to cross a system boundary.\n- [AG-UI](https://github.com/ag-ui-protocol/ag-ui) — Lightweight event-driven protocol standardizing how AI agents connect to frontend applications: streaming state updates, tool call rendering, and HITL interrupts over a shared event bus. Fills the layer between MCP (tool access) and A2A (agent-to-agent) — it's the missing protocol for real-time agent-to-UI communication that neither MCP nor A2A was designed to address. ![Stars](https://img.shields.io/github/stars/ag-ui-protocol/ag-ui?style=flat-square\u0026label=★\u0026color=yellow)\n- [Code Execution with MCP: Building More Efficient Agents](https://www.anthropic.com/engineering/code-execution-with-mcp) — Anthropic's engineering account of reducing tool-call token overhead by having agents write code to interact with MCP servers rather than calling tools directly: up to 98.7% token reduction in experiments. Broadly applicable to any harness where tool schema overhead and intermediate results are consuming context — the pattern is to wrap multi-step tool interaction in a code execution primitive rather than exposing each operation as a discrete tool call.\n- [Microsoft Skills Framework](https://github.com/microsoft/skills) — Standardized framework for defining, versioning, and distributing agent skills. Enables skill reuse across Claude Code, Copilot, VS Code, Gemini, and other platforms — a harness-level abstraction that makes skills first-class deployment artifacts rather than ad-hoc tool definitions. ![Stars](https://img.shields.io/github/stars/microsoft/skills?style=flat-square\u0026label=★\u0026color=yellow)\n- [SkillNet \u0026 SkillsBench: Infrastructure for AI Agent Skills at Scale](https://github.com/skillmatic-ai/awesome-agent-skills) — Comprehensive framework for creating, evaluating, and sharing agent skills with 86-task benchmark across 11 domains. Demonstrates the harness problem of skill fragmentation and provides infrastructure for standardized skill evaluation across frameworks. ![Stars](https://img.shields.io/github/stars/skillmatic-ai/awesome-agent-skills?style=flat-square\u0026label=★\u0026color=yellow)\n- [Comet](https://github.com/rpamis/comet) — Resumable long-running task workflow and Skill platform for coding that turns skill creation, evaluation, and release into a single lifecycle with Rubric, Pass@k, and Pass^k scoring. The Native/Classic dual-workflow model is a concrete reference for matching harness constraint strength to model capability rather than using one loop for every task. ![Stars](https://img.shields.io/github/stars/rpamis/comet?style=flat-square\u0026label=★\u0026color=yellow)\n- [addyosmani/agent-skills](https://github.com/addyosmani/agent-skills) — Production-grade engineering skills for AI coding agents, packaged as 24 reusable skills covering the full development lifecycle from `/spec` to `/ship`. The slash-command interface and context-aware auto-activation make it a concrete reference for turning senior-engineering judgment into agent-executable harness artifacts. ![Stars](https://img.shields.io/github/stars/addyosmani/agent-skills?style=flat-square\u0026label=★\u0026color=yellow)\n- [AWS Bedrock AgentCore with WebRTC Support](https://aws.amazon.com/about-aws/whats-new/2026/03/amazon-bedrock-webrtc/) — Adds peer-to-peer, UDP-based WebRTC bidirectional streaming to Bedrock Agents for real-time voice interactions. Complements existing WebSocket support with lower latency and better resilience for poor network conditions. Essential harness-level transport choice for agents targeting sub-800ms Total Turn-Around Time voice interactions.\n- [Hermes Agent: Unified Streaming for Real-Time Agent Workflows](https://juliangoldie.com/hermes-agent-unified-streaming/) — Token-by-token streaming delivery system enabling real-time agent responses; sub-second decision loops on streaming events vs. batch-refreshed data. Critical infrastructure for harnesses where latency (not just throughput) is the constraint — agents must react to events as they arrive, not wait for batch completions.\n- [Google Developers: Closing the Knowledge Gap with Agent Skills](https://developers.googleblog.com/closing-the-knowledge-gap-with-agent-skills/) — Google ADK expansion with evaluation harness (117 prompts) for assessing skill performance across agentic coding, chatbots, document processing. Provides reference patterns and benchmark datasets for skill evaluation, complementing the Microsoft Skills Framework with Google's evaluation methodology.\n- [What's New with GitHub Copilot Coding Agent](https://github.blog/ai-and-ml/github-copilot/whats-new-with-github-copilot-coding-agent/) — GitHub's February 26, 2026 update is worth including for one specific reason: it makes `.github/agents/` custom agent files, self-review, built-in security scanning, and CLI handoff concrete as harness primitives rather than abstract ideas. Useful as a current reference for how repository-scoped agent definitions and security checks are being productized in a real coding-agent control plane.\n- [Announcing Official MCP Support for Google Services](https://cloud.google.com/blog/products/ai-machine-learning/announcing-official-mcp-support-for-google-services) — Google's 2026 rollout of managed MCP endpoints is a useful counterpoint to self-hosted MCP servers: discovery, IAM, audit logging, and Model Armor are provided as platform primitives instead of being rebuilt per server. Worth including because it shows what \"enterprise MCP\" looks like when the transport, auth, and governance layers are treated as product surface rather than glue code.\n- [Dataverse Skills: Your Coding Agent Now Speaks Dataverse](https://devblogs.microsoft.com/powerplatform/dataverse-skills-your-coding-agent-now-speaks-dataverse) — Microsoft's April 1, 2026 release is a strong concrete example of domain-specific skills done properly: the agent learns when to use MCP, when to drop to a Python SDK, and when to call a raw API, while the user stays in natural language. Worth adding because it shows that \"skills\" are not just prompt snippets, but curated execution strategies that hide a multi-tool integration stack behind intent.\n- [You can't whisper at an AI agent](https://stripe.dev/blog/ai-steering-experiments) — Stripe's May 2026 study of how agents actually consume SDK and CLI guidance: passive documentation is ignored, while hard steering signals placed in the loaded context — skill files, error messages, CLI prompts — reliably change behavior. The \"if your guidance wasn't in the loaded context, it didn't happen\" principle is a practical design rule for any skill, tool, or SDK that agents are expected to use.\n- [Agent Toolkit for AWS](https://github.com/aws/agent-toolkit-for-aws) — Official AWS-supported MCP servers, skills, and plugins that let AI agents provision, query, and manage AWS resources through a standardized protocol interface. Worth including as the reference for how a major cloud provider productizes infrastructure access into agent-ready harness primitives rather than leaving teams to hand-roll IAM-scoped tool definitions. ![Stars](https://img.shields.io/github/stars/aws/agent-toolkit-for-aws?style=flat-square\u0026label=★\u0026color=yellow)\n- [agentic-stack](https://github.com/codejunkie99/agentic-stack) — Portable `.agent/` folder that externalizes memory, skills, and protocols from any specific coding agent into a cross-tool harness layer. Adapters translate the same configuration into Claude Code's `CLAUDE.md`, Cursor's rules, OpenCode's `AGENTS.md`, and more — the first practical answer to harness vendor lock-in at the configuration level. ![Stars](https://img.shields.io/github/stars/codejunkie99/agentic-stack?style=flat-square\u0026label=★\u0026color=yellow)\n- [mcp-agent](https://github.com/lastmile-ai/mcp-agent) — Production-grade framework for building agents with MCP: composable workflows, built-in observability, and provider-agnostic model routing. The clearest reference for turning MCP servers from isolated utilities into a coherent agent harness. ![Stars](https://img.shields.io/github/stars/lastmile-ai/mcp-agent?style=flat-square\u0026label=★\u0026color=yellow)\n- [vurb.ts](https://github.com/vinkius-labs/vurb.ts) — TypeScript framework for building production MCP servers with a \"Presenter\" perception layer that strips undeclared fields, redacts PII, and gates tool visibility by workflow state. Fills a critical gap in the MCP ecosystem: most tooling focuses on consuming servers, while vurb.ts addresses the harness engineering of authoring servers that are safe, governable, and context-aware by default. ![Stars](https://img.shields.io/github/stars/vinkius-labs/vurb.ts?style=flat-square\u0026label=★\u0026color=yellow)\n- [SkillOpt](https://github.com/microsoft/SkillOpt) — Microsoft's skill optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits and validation-gated updates, producing deployable `best_skill.md` artifacts. The key harness insight is that skills should be treated as optimizable parameters that improve with execution feedback, not static prompt fragments written once and forgotten. ![Stars](https://img.shields.io/github/stars/microsoft/SkillOpt?style=flat-square\u0026label=★\u0026color=yellow)\n- [superpowers](https://github.com/obra/superpowers) — Agentic skills framework and software-development methodology with automatically-triggered, mandatory skills that work across Claude Code, Cursor, Codex, Gemini CLI, and Copilot CLI. Demonstrates how to package cross-harness workflows — TDD, subagent-driven development, review gates — as reusable skills with an eval harness. ![Stars](https://img.shields.io/github/stars/obra/superpowers?style=flat-square\u0026label=★\u0026color=yellow)\n- [Antigravity Awesome Skills](https://github.com/sickn33/antigravity-awesome-skills) — Installable library of 1,400+ agentic skills for Claude Code, Cursor, Codex CLI, Gemini CLI, and more. The largest community-driven skill catalog with an npm installer and role-based bundles — a concrete reference for treating skills as versioned harness artifacts rather than ad-hoc prompts. ![Stars](https://img.shields.io/github/stars/sickn33/antigravity-awesome-skills?style=flat-square\u0026label=★\u0026color=yellow)\n- [agentgateway](https://github.com/agentgateway/agentgateway) — Open-source agentic proxy that unifies LLM gateway, MCP gateway, and A2A gateway into a single control plane for managing multi-agent, multi-tool connectivity at scale. Provides drop-in security, observability, and governance for agent-to-LLM, agent-to-tool, and agent-to-agent communication — the missing infrastructure layer between agents and the services they touch. ![Stars](https://img.shields.io/github/stars/agentgateway/agentgateway?style=flat-square\u0026label=★\u0026color=yellow)\n- [AIP: A Graph Representation for Learning and Governing Agent Skills](https://arxiv.org/abs/2606.04781) — June 2026 proposal to replace free-form skill prose with directed execution graphs: discrete steps as nodes backed by deterministic scripts or natural-language descriptions, connected by explicit typed input/output edges and governed by a schema-validated YAML spec. Compiling skills to AIP improved Claude Sonnet's pass rate from 53% to 67% while making skills queryable, auditable, and repairable at the script level — the key harness advance is turning skill improvement from a prose rewrite into a measurable tuning loop.\n- [Ponytail](https://github.com/DietrichGebert/ponytail) — Skill that makes coding agents behave like a \"lazy senior dev\": prefer built-in solutions, avoid new dependencies, and write the minimum code that works. Benchmarked on real Claude Code sessions with ~54% fewer lines, ~20% lower cost, and preserved safety guards — a rare harness-level incentive that fights over-engineering rather than just adding capability. ![Stars](https://img.shields.io/github/stars/DietrichGebert/ponytail?style=flat-square\u0026label=★\u0026color=yellow)\n- [wshobson/agents](https://github.com/wshobson/agents) — Cross-harness plugin marketplace that maintains one source-of-truth `plugins/` directory and generates harness-native artifacts for Claude Code, Codex CLI, Cursor, OpenCode, Gemini CLI, and GitHub Copilot. It is the clearest practical example of treating reusable agent capabilities as a portable distribution format rather than ad-hoc prompt files. ![Stars](https://img.shields.io/github/stars/wshobson/agents?style=flat-square\u0026label=★\u0026color=yellow)\n- [mgechev/skillgrade](https://github.com/mgechev/skillgrade) — CLI that turns agent skill verification into repeatable unit tests: it generates task and grader pairs from a `SKILL.md`, runs multi-trial evals against Claude, Codex, Gemini, or OpenCode, and reports pass rates with a CI-ready threshold. Fills the gap between shipping a skill and knowing an agent actually discovers and invokes it correctly. ![Stars](https://img.shields.io/github/stars/mgechev/skillgrade?style=flat-square\u0026label=★\u0026color=yellow)\n- [Qwen-MM-Plugins](https://github.com/QwenLM/Qwen-MM-Plugins) — Qwen's official multimodal plugin suite packages vision, video, document, 3D, and CAD capabilities as portable skills and MCP servers across Claude Code, Codex, OpenCode, and other harnesses. It shows how to make a text-first coding harness multimodal-native without rebuilding the agent loop. ![Stars](https://img.shields.io/github/stars/QwenLM/Qwen-MM-Plugins?style=flat-square\u0026label=★\u0026color=yellow)\n\n### Permissions \u0026 Authorization\n\n- [Beyond Permission Prompts](https://www.anthropic.com/engineering/beyond-permission-prompts) — Structured authorization patterns for agents: how to give agents the right permissions without relying on prompt-level trust.\n- [OWASP LLM06:2025 — Excessive Agency](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/) — OWASP's authoritative definition of the \"excessive agency\" risk: over-provisioned functions, unnecessary permissions, and missing approval mechanisms. The standard checklist for auditing harness permission scope against principle of least privilege.\n- [GitHub Enterprise — Governing Agents](https://wellarchitected.github.com/library/governance/recommendations/governing-agents/) — April 2026 GitHub official guide for enterprise agent governance: MCP server registry curation with ruleset-protected configurations, agent environment standardization via `copilot-setup-steps.yml`, ephemeral runner enforcement, and cloud-agent firewall allowlisting. The most concrete published reference for governing agent fleets at scale without creating bottlenecks.\n- [Claude Code Auto Mode: A Safer Way to Skip Permissions](https://www.anthropic.com/engineering/claude-code-auto-mode) — Anthropic's engineering post on replacing approval fatigue (users approve 93% of prompts, making approvals meaningless) with a two-stage classifier: fast single-token gate first, chain-of-thought reasoning only on flagged actions. The design decisions — stripping assistant messages to prevent the agent from rationalizing dangerous actions, deny-and-continue recovery instead of halt — are the reference design for safe-by-default headless agent permissions.\n- [Claude Agent SDK — Configure Permissions](https://platform.claude.com/docs/en/agent-sdk/permissions) — The most concrete reference for harness permission architecture: five-layer evaluation order (hooks → deny rules → permission mode → allow rules → canUseTool), `allowedTools`/`disallowedTools` declarative scoping, and four permission modes including `dontAsk` (deny-by-default for headless agents). The subagent inheritance warning for `bypassPermissions` alone is worth reading before any multi-agent deployment.\n- [Two Different Types of Agent Authorization](https://blog.langchain.com/two-different-types-of-agent-authorization/) — Distinguishes on-behalf-of authorization (agent uses end-user credentials, requires cross-channel identity mapping and per-user memory isolation) from fixed-credential authorization (agent owns its own account, requires human-in-the-loop guardrails on high-risk actions). The two models have fundamentally different threat surfaces and determine where authorization enforcement lives in the harness.\n- [Authorization and Governance for AI Agents: Runtime Authorization Beyond Identity at Scale](https://techcommunity.microsoft.com/blog/microsoft-security-blog/authorization-and-governance-for-ai-agents-runtime-authorization-beyond-identity/4509161) — Microsoft Security's reusable Authorization Fabric combining a Policy Enforcement Point (PEP) and Policy Decision Point (PDP) as a Microsoft Entra-protected endpoint. Every agent calls this fabric before tool execution, receiving a deterministic decision: ALLOW / DENY / REQUIRE_APPROVAL / MASK. Addresses the gap that identity alone (who is this agent?) doesn't answer whether a specific action should be executed now, by this agent, for this user, under the current business and regulatory context.\n- [IETF draft-klrc-aiagent-auth: AI Agent Authentication and Authorization](https://datatracker.ietf.org/doc/draft-klrc-aiagent-auth/) — The first IETF standards-track specification for AI agent authentication (March 2026, authors from AWS, OpenAI, Zscaler, Ping Identity, Defakto Security). Builds on WIMSE (Workload Identity in Multi-System Environments) and OAuth 2.0 rather than inventing new protocols — agents get SPIFFE-style identifiers, with delegation via OAuth Token Exchange and DPoP for token binding. Essential reference for any harness that needs to authenticate agents across trust domains.\n- [Nango: Pre-Built Authentication for AI Agents](https://nango.dev) — Open-source platform providing pre-built OAuth and API key authentication for 700+ APIs across 30 categories. Automatically refreshes access tokens, provides webhooks when credentials break, and stores tokens securely so agent code never touches secrets. Solves the \"agent needs to call an authenticated API\" problem at scale — the authentication layer that complements Composio's tool wrapping. ![Stars](https://img.shields.io/github/stars/NangoHQ/nango?style=flat-square\u0026label=★\u0026color=yellow)\n- [Agent Vault](https://github.com/Infisical/agent-vault) — Infisical's open-source credential broker that sits between AI agents and the APIs they call, injecting real credentials onto outbound requests so agents never possess secrets directly. Eliminates a concrete prompt-injection attack surface — exfiltration of API keys and PATs — by treating credential possession as a harness-layer boundary rather than an agent-side configuration. ![Stars](https://img.shields.io/github/stars/Infisical/agent-vault?style=flat-square\u0026label=★\u0026color=yellow)\n- [AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security](https://arxiv.org/abs/2601.18491) — Three-dimensional risk taxonomy (source/failure-mode/consequence) with fine-grained agentic safety benchmark (ATBench) and diagnostic guardrail models (4B–8B parameters) achieving 91.8% accuracy. Shifts safety monitoring from binary safe/unsafe checks to root-cause diagnosis: why did an action violate constraints? Where did the violation originate? What are the downstream consequences? Essential for production harnesses where transparency into safety decisions is required for audit trails.\n- [Open Agent Passport (OAP): Deterministic Pre-Action Authorization for Autonomous AI Agents](https://arxiv.org/abs/2603.20953) — March 2026 open specification and reference implementation that intercepts tool calls synchronously before execution, evaluates them against a declarative policy, and produces a cryptographically signed audit record. Enforces authorization in a median of 53ms; in a live adversarial testbed ($5,000 bounty), restrictive OAP policy achieved 0% attack success vs. 74.6% under permissive policies. Distinguishes pre-action authorization from sandboxed execution and model-based screening as complementary but distinct harness layers.\n- [nah](https://github.com/manuelschipper/nah) — Deterministic permission guard that maps tool calls to an intent taxonomy (`filesystem_delete`, `network_outbound`, `lang_exec`, etc.) rather than relying on command-name allow/deny lists. The key insight for harness design: the same binary can be benign or destructive depending on its arguments, so intent-level enforcement is the only reproducible safety layer. ![Stars](https://img.shields.io/github/stars/manuelschipper/nah?style=flat-square\u0026label=★\u0026color=yellow)\n- [Aegis](https://github.com/Justin0504/Aegis) — Pre-execution firewall that intercepts, classifies, and blocks agent tool calls before they execute, with a compliance cockpit for real-time monitoring, human-in-the-loop approvals, and a tamper-evident audit trail. The zero-code-change integration makes runtime policy enforcement practical for existing agent deployments. ![Stars](https://img.shields.io/github/stars/Justin0504/Aegis?style=flat-square\u0026label=★\u0026color=yellow)\n- [When \"Do Not\" Is Not Deny: Security Rules in CLAUDE.md vs Built-In Controls](https://arxiv.org/abs/2608.23550) — Analyzes 481 public CLAUDE.md files and finds only ~4% of natural-language security rules are backed by a matching built-in control, exposing the gap between documented intent and enforced permissions. A concrete reminder that harness instructions are not guardrails unless they map to deterministic enforcement.\n- [Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?](https://arxiv.org/abs/2608.27443) — August 2026 user study (113 non-technical participants) comparing per-action HITL approval, automated model review, and user-authored allow/ask/never rules. The key harness finding: pre-authored policies blocked ~20 percentage points *less* overreach than per-action approval — users overwhelmingly set \"ask\" as their rule and then approve anyway, so static rule pre-commitment degenerates back into runtime prompts without gaining protection. Essential evidence that permission UX design (not just policy expressiveness) determines whether stated intent survives contact with real agent actions.\n\n### Memory \u0026 State\n\n- [Building Effective Agents](https://www.anthropic.com/research/building-effective-agents) — Covers in-context, external, and procedural memory patterns as harness-level concerns.\n- [Letta (MemGPT)](https://github.com/letta-ai/letta) — The reference architecture for stateful agents: three-tier memory (core / archival / recall) maps directly to harness state management design. Their [agent loop redesign post](https://www.letta.com/blog/letta-v1-agent) is the most thorough public analysis of how memory structure shapes the harness. ![Stars](https://img.shields.io/github/stars/letta-ai/letta?style=flat-square\u0026label=★\u0026color=yellow)\n- [mem0](https://github.com/mem0ai/mem0) — Drop-in universal memory layer (YC-backed, AWS Agent SDK's exclusive memory provider) that handles cross-session retention without custom harness-level state management code. Lowest integration cost for production-grade persistent memory. ![Stars](https://img.shields.io/github/stars/mem0ai/mem0?style=flat-square\u0026label=★\u0026color=yellow)\n- [Stash](https://github.com/alash3al/stash) — Self-hosted persistent memory layer with an 8-stage consolidation pipeline (episodes → facts → relationships → patterns) and built-in MCP server. The critical gap it fills: production-grade cross-session memory without cloud dependencies or complex infrastructure — a single Docker Compose gives you Postgres, pgvector, and background consolidation. ![Stars](https://img.shields.io/github/stars/alash3al/stash?style=flat-square\u0026label=★\u0026color=yellow)\n- [TencentDB-Agent-Memory](https://github.com/Tencent/TencentDB-Agent-Memory) — Tencent's fully local agent memory system with a 4-tier progressive pipeline (Conversation → Atom → Scenario → Persona) and symbolic short-term memory via Mermaid canvases. The benchmark data is striking: 61% token reduction and 51% relative pass-rate improvement on long-horizon tasks, demonstrating that hierarchical memory architecture outperforms flat vector stores for coding agents. ![Stars](https://img.shields.io/github/stars/Tencent/TencentDB-Agent-Memory?style=flat-square\u0026label=★\u0026color=yellow)\n- [Zep](https://github.com/getzep/zep) — Purpose-built agent memory store with automatic conversation summarization, entity extraction, and semantic search over session history. Solves long-session context overflow at the memory layer rather than forcing the harness to manage trimming manually. ![Stars](https://img.shields.io/github/stars/getzep/zep?style=flat-square\u0026label=★\u0026color=yellow)\n- [engram](https://github.com/Gentleman-Programming/engram) — Persistent memory for AI coding agents delivered as a single Go binary with SQLite + FTS5 and 18 MCP tools for save, search, session lifecycle, and conflict detection. The agent-agnostic, zero-dependency design makes cross-session memory a local harness primitive rather than a managed cloud service. ![Stars](https://img.shields.io/github/stars/Gentleman-Programming/engram?style=flat-square\u0026label=★\u0026color=yellow)\n- [MemPalace](https://github.com/MemPalace/mempalace) — Local-first AI memory system that stores conversation history verbatim and retrieves it with semantic search through a structured palace architecture (wings, rooms, drawers). Achieves 96.6% R@5 on LongMemEval with zero LLM calls, making it the best-benchmarked open-source memory layer for agents that need cross-session persistence without cloud dependencies. ![Stars](https://img.shields.io/github/stars/MemPalace/mempalace?style=flat-square\u0026label=★\u0026color=yellow)\n- [agentmemory](https://github.com/rohitg00/agentmemory) — Persistent memory layer purpose-built for coding agents with 95.2% retrieval accuracy and 92% token reduction, backed by real-world benchmarks. Its cross-agent architecture — one memory server serving Claude Code, Cursor, Codex, and OpenCode through MCP and hooks — makes cross-session memory a portable harness primitive rather than a vendor-specific add-on. ![Stars](https://img.shields.io/github/stars/rohitg00/agentmemory?style=flat-square\u0026label=★\u0026color=yellow)\n- [claude-memory-compiler](https://github.com/coleam00/claude-memory-compiler) — Turns raw Claude Code sessions into a self-evolving knowledge base: hooks capture every interaction, the Agent SDK extracts decisions and lessons, and an LLM compiler distills them into structured, cross-referenced articles that improve retrieval quality over time. The most concrete open-source implementation of trace-driven memory evolution for coding agents. ![Stars](https://img.shields.io/github/stars/coleam00/claude-memory-compiler?style=flat-square\u0026label=★\u0026color=yellow)\n- [mex](https://github.com/mex-memory/mex) — Turns agent-learned project knowledge into a repo-local, symbol-grounded wiki with drift detection and task-aware routing. The critical gap it fills: most agent memory stores facts, but mex keeps those facts connected to the exact code symbols they describe and flags when code changes invalidate them — making memory freshness a harness-level concern rather than a manual cleanup task. ![Stars](https://img.shields.io/github/stars/mex-memory/mex?style=flat-square\u0026label=★\u0026color=yellow)\n- [How We Built Agent Builder's Memory System](https://blog.langchain.com/how-we-built-agent-builders-memory-system/) — LangChain's engineering account of a COALA-based three-tier memory system (procedural/semantic/episodic) backed by PostgreSQL but exposed to agents as a virtual filesystem. Key harness decisions: human-in-the-loop approval gates every memory write (blocking prompt-injection via malformed writes), validation errors are fed back to the LLM for self-correction, and AGENTS.md serves as the agent's procedural memory anchor.\n- [Building an Agentic Memory System for GitHub Copilot](https://github.blog/ai-and-ml/github-copilot/building-an-agentic-memory-system-for-github-copilot/) — GitHub's January 15, 2026 write-up is one of the clearest public discussions of deployed cross-agent memory: repository-scoped memories are shared across coding agent, CLI, and code review, but only after just-in-time verification against the current code state. The core harness lesson is that memory quality is mostly a freshness and invalidation problem — stale, branch-specific memories are often more dangerous than having no memory at all.\n- [Building Self-Correcting Memory in OpenWiki](https://www.langchain.com/blog/self-correcting-memory-openwiki) — LangChain's August 2026 design account of memory that can forget: every claim the agent writes is persisted with versioned code evidence, and when that evidence changes the claim is flagged stale — remaining durably uncertain until re-verified rather than silently trusted. The most concrete published mechanism for memory invalidation: staleness becomes a first-class, inspectable harness state instead of a side effect of periodic regeneration.\n- [MemArchitect: A Policy-Driven Memory Governance Layer](https://arxiv.org/abs/2603.18330) — Proposes a governance layer that decouples memory lifecycle management (decay, conflict resolution, privacy enforcement) from model weights, directly addressing the \"zombie memory\" problem: outdated facts sitting in the context window that only a harness-level eviction policy — not the model — can remove.\n- [Codified Context: Infrastructure for AI Agents in a Complex Codebase](https://arxiv.org/abs/2602.20478) — Production-validated architecture (283 sessions, 108k-line codebase) built on three components: a \"hot-memory constitution\" encoding conventions and multi-agent coordination protocols, 19 domain-specialist agents, and a \"cold-memory knowledge base\" of 34 on-demand specification documents. The empirical data distinguishes what must live in always-on context from what should be retrieved on demand — the most concrete published guidance for scaling cross-session memory in a large codebase.\n- [Facts as First Class Objects: Knowledge Objects for Persistent LLM Memory](https://arxiv.org/abs/2603.17781) — Identifies three production failure modes of in-context memory at scale: capacity overflow at ~8,000 facts, 60% fact destruction during compaction, and 54% behavioral drift from constraint erosion across cascaded summarizations. Proposes Knowledge Objects (hash-addressed discrete fact tuples) achieving 100% accuracy at 252× lower cost than in-context storage — the quantitative case for moving persistent facts out of the context window into a structured retrieval layer rather than managing them through prompt engineering.\n- [Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents](https://arxiv.org/abs/2601.22352) — Formal framework for measuring how well agents recover from tool failures. Defines Expected Recovery Regret (ERR) as a metric for harness design: the cost of recovering from stochastic failures in downstream tasks. Critical for assessing reliability of production harnesses where tool calls occasionally fail but agents must continue functioning.\n- [MAGMA: Multi-Graph Agentic Memory Architecture](https://arxiv.org/abs/2601.03236) — Represents agent memory across four orthogonal semantic, temporal, causal, and entity graphs, enabling policy-guided retrieval over relational views. Outperforms MemGPT on long-horizon reasoning benchmarks by 18.5% accuracy improvement. The multi-graph abstraction lets harness engineers compose different retrieval strategies for different task phases — a concrete architecture for memory that scales beyond single-view approaches.\n- [GAAMA: Graph Augmented Associative Memory for Agents](https://arxiv.org/abs/2603.27910) — Hybrid memory system blending graph traversal with semantic similarity through additive scoring; graph augmentation improves retrieval over embedding-only approaches for long-horizon reasoning. Practical alternative to full multi-graph systems when adding structure to existing vector-based memory is sufficient.\n- [Graph-Native Cognitive Memory for AI Agents: Formal Belief Revision Semantics for Versioned Memory Architectures](https://arxiv.org/abs/2603.17244) — Formal semantics for versioned memory graphs with belief revision operations, enabling agents to maintain coherent evolving world models through multi-turn reasoning. Addresses the hard problem of inconsistency resolution in long-lived agent memory: when new information contradicts prior beliefs, how should the agent update its knowledge base?\n- [Continual learning for AI agents](https://blog.langchain.com/continual-learning-for-ai-agents/) — LangChain's April 2026 framing of agent learning as three distinct layers: model weights, harness behavior, and contextual memory. Essential for designing memory systems that don't just store facts but actually improve agent performance over time through trace-driven harness and context updates.\n- [cognee](https://github.com/topoteretes/cognee) — Open-source memory platform with a hybrid graph-vector-relational poly-store that lets agents recall facts through both semantic similarity and structured graph traversal — the practical middle ground between flat vector stores and full multi-graph research systems. The self-improving pipeline reweights edges from agent feedback, and the 14-tool MCP server makes it a drop-in memory primitive for production harnesses. ![Stars](https://img.shields.io/github/stars/topoteretes/cognee?style=flat-square\u0026label=★\u0026color=yellow)\n- [Hindsight](https://github.com/vectorize-io/hindsight) — Agent memory system organized around three explicit operations—retain, recall, and reflect—with semantic, keyword, graph, and temporal retrieval plus an MCP server. The June 2026 release and production usage make it a concrete reference for turning cross-session persistence from passive storage into an active learning layer inside the harness. ![Stars](https://img.shields.io/github/stars/vectorize-io/hindsight?style=flat-square\u0026label=★\u0026color=yellow)\n- [ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents](https://arxiv.org/abs/2604.10352) — Applies virtual-memory semantics to agent context management, treating the context window as working memory with typed pages, minimum-fidelity invariants, and validated writeback at lifecycle boundaries. A concrete reference for making residency and durability auditable harness-level concerns rather than best-effort side effects.\n- [MAGE: Memory as Agent-Guided Exploration](https://arxiv.org/abs/2606.06090) — June 2026 proposal to treat long-horizon memory as execution-state management rather than semantic retrieval: a hierarchical state tree preserves trajectories, enables rollback, and constructs working state from the active root-to-leaf path. Improves task success by 7.8–20.4 percentage points while cutting token consumption 55.1%, making it a concrete reference for memory architectures where execution history — not just semantic similarity — determines recall.\n- [deja-vu](https://github.com/vshulcz/deja-vu) — Indexes coding-agent sessions already written to disk and serves them back over MCP with no LLM calls, embeddings, or API keys. The zero-dependency binary solves the memory cold-start problem for cross-session persistence and demonstrates that agent memory can be built on existing filesystem traces rather than a separate learning pipeline. ![Stars](https://img.shields.io/github/stars/vshulcz/deja-vu?style=flat-square\u0026label=★\u0026color=yellow)\n- [trajectory](https://github.com/letta-ai/trajectory) — Letta's July 2026 library that normalizes native session transcripts from 15+ harnesses (Claude Code, Codex, Cursor, OpenCode, Pi, Gemini CLI, OpenHands, and more) into one validated, model-ready record format. Fills a real infrastructure gap — every runtime logs the same concepts (messages, reasoning, tool calls, results) in incompatible formats — and is explicitly designed for *agent* consumption (memory formation, search, training), turning raw session logs into a portable cross-harness substrate rather than yet another observability dashboard. ![Stars](https://img.shields.io/github/stars/letta-ai/trajectory?style=flat-square\u0026l","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/ai-boost%2Fawesome-harness-engineering/projects"}