Projects in Awesome Lists tagged with agent-reliability
A curated list of projects in awesome lists tagged with agent-reliability .
https://github.com/qualixar/superlocalmemory
World's first local-only AI memory to break 74% retrieval and 60% zero-LLM on LoCoMo. No cloud, no APIs, no data leaves your machine. Additionally, mode C (LLM/Cloud) - 87.7% LoCoMo. Research-backed. arXiv: 2603.14588
agent-memory agent-reliability ai-agents claude-code cursor knowledge-graph llm-memory local-first mcp mcp-server persistent-memory qualixar semantic-search vector-search windsurf
Last synced: 03 Jun 2026
https://github.com/IgorGanapolsky/ThumbGate
Agent governance for ThumbGate: 👍/👎 become Pre-Action Checks that block repeat mistakes before code, money, or customer systems change.
agent-reliability ai-agents ai-cost-optimization ai-safety amp claude-code codex cursor developer-tools feedback-loop gemini guardrails mcp mcp-server opencode pre-action-checks reduce-llm-cost save-llm-tokens thompson-sampling thumbgate
Last synced: 12 Jun 2026
https://github.com/igorganapolsky/thumbgate
Agent governance for ThumbGate: 👍/👎 become Pre-Action Checks that block repeat mistakes before code, money, or customer systems change.
agent-reliability ai-agents ai-cost-optimization ai-safety amp claude-code codex cursor developer-tools feedback-loop gemini guardrails mcp mcp-server opencode pre-action-checks reduce-llm-cost save-llm-tokens thompson-sampling thumbgate
Last synced: 30 May 2026
https://github.com/yzhao062/auditable
Audit any agent decision across its past, present, and future, on one typed graph.
agent-reliability agentic-ai ai-agents ai-risk ai-safety anomaly-detection llm llm-agents llmops observability python trustworthy-ai
Last synced: 26 Jun 2026
https://github.com/ek33450505/attest
DONE is a claim, not proof. A local, deterministic, zero-LLM Claude Code hook that verifies a subagent's Status: DONE / ## Handoff claim against the real git working-tree delta — and, opt-in, blocks a DONE whose claimed files never actually landed on disk. It adds no tokens, cannot itself hallucinate, and fails open on every doubt.
agent-reliability claude-code claude-code-hooks developer-tools llm-agents
Last synced: 30 Jun 2026
https://github.com/hugomn/mast-taxonomy-production-telemetry
Applying the MAST failure taxonomy (Cemri et al. 2025) to 639k execution steps from a production autonomous-agent platform. Honest, reproducible failure-mode analysis.
agent-reliability ai-agents failure-analysis llm-agents llm-evaluation mast observability production-ml
Last synced: 15 Jul 2026