{"id":128155,"url":"https://github.com/Eric-LLMs/Awesome-AI-Engineering","name":"Awesome-AI-Engineering","description":"A full-stack LLM engineering playbook- practical guides for building, deploying, and evaluating LLM systems and AI agents.","projects_count":71,"last_synced_at":"2026-09-02T18:00:21.077Z","repository":{"id":329290346,"uuid":"1118956298","full_name":"Eric-LLMs/Awesome-AI-Engineering","owner":"Eric-LLMs","description":"A full-stack LLM engineering playbook- practical guides for building, deploying, and evaluating LLM systems and AI agents.","archived":false,"fork":false,"pushed_at":"2026-08-06T06:41:35.000Z","size":60120,"stargazers_count":13,"open_issues_count":0,"forks_count":1,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-08-14T05:26:29.158Z","etag":null,"topics":["agents","ai-infrastructure","deep-learning","evaluation","evaluation-framework","fine-tuning","langchain","llm","machine-learning","mcp","mcp-server","nlp","prompt","rag","security","system-design"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Eric-LLMs.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"claude":null,"gemini":null,"cursor":null,"copilot":null,"dco":null,"cla":null,"disclosure":null}},"created_at":"2025-12-18T14:26:45.000Z","updated_at":"2026-08-06T06:41:39.000Z","dependencies_parsed_at":"2026-08-14T01:27:41.844Z","dependency_job_id":null,"html_url":"https://github.com/Eric-LLMs/Awesome-AI-Engineering","commit_stats":null,"previous_names":["eric-llms/booknotes","eric-llms/ai-engineering-notes","eric-llms/awesome-ai-engineering"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/Eric-LLMs/Awesome-AI-Engineering","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Eric-LLMs%2FAwesome-AI-Engineering","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Eric-LLMs%2FAwesome-AI-Engineering/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Eric-LLMs%2FAwesome-AI-Engineering/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Eric-LLMs%2FAwesome-AI-Engineering/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Eric-LLMs","download_url":"https://codeload.github.com/Eric-LLMs/Awesome-AI-Engineering/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Eric-LLMs%2FAwesome-AI-Engineering/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":37019084,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-02T02:00:06.194Z","response_time":106,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2026-04-14T06:01:00.324Z","updated_at":"2026-09-02T18:00:21.078Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["📑 Table of Contents"],"sub_categories":["🧰 Key Tools, Frameworks \u0026 Strategies","🛠️ Hands-on Projects \u0026 Tools","🧰 Key Frameworks \u0026 Code Samples","🧰 Key Frameworks \u0026 Tools","🧰 Key Open-Source Projects \u0026 References","🔗 Related Protocols","📑 Further Reading / Resources","🛠️ Hands-on: A Minimal ReAct Agent","🛠️ Hands-on Lab \u0026 Examples","🔑 Mind Map (Key Concepts)","📑 Presentation Slides","🔑 Mind Map (Framework Overview)"],"readme":"# 📘 Awesome AI Engineering   \n\n[![License: CC BY 4.0](https://img.shields.io/badge/License-CC_BY_4.0-lightgrey.svg)](https://creativecommons.org/licenses/by/4.0/)\n\n### The Full-Stack LLM Engineering Playbook\n\nA full-stack LLM engineering playbook — practical guides for building, deploying, and evaluating LLM systems and AI agents. For a deep dive into the research frontier, explore the [📖 LLM Technology Landscape \u0026 Evolution](https://github.com/Eric-LLMs/LLMs-Lab/tree/main/Docs) — a curated reading list covering the full LLM stack, from model architectures and training, fine-tuning, inference optimization, reasoning, and Agent systems.  \n\n\u003ca id=\"top\"\u003e\u003c/a\u003e\n\n## 📑 Table of Contents\n\n| 📚 Content                                                       | 🔗 Quick Link                                                 |\n|:----------------------------------------------------------------------|:--------------------------------------------------------------|\n| Building LLMs for Production                                          | [🔍 Explore](#building-llms-for-production)                   |\n| Building High-Performance, Private AI Infrastructure for the Enterprise                          | [🔍 Explore](#high-performance-private-ai-infrastructure)           |\n| Building AI Agents                                             | [🔍 Explore](#building-ai-agents)                      |\n| Mastering the Model Context Protocol (MCP)                            | [🔍 Explore](#mastering-the-model-context-protocol)           |\n| Agent Memory Part I  (A Survey of Memory)                             | [🔍 Explore](#agent-memory-part-i)                            |\n| Agent Memory Part II (Building Memory Modules for Agentic AI Systems) | [🔍 Explore](#building-memory-modules-for-agentic-ai-Systems) |\n| Agent Evaluation (Eval) Engineering                                   | [🔍 Explore](#agent-eval)                        |\n---\n\u003cbr\u003e  \n\u003ca id=\"building-llms-for-production\"\u003e\u003c/a\u003e   \n \n# 📚 Building LLMs for Production \nThis guide covers LLM production, from Transformer architectures to advanced techniques like RAG and Fine-Tuning. It explores frameworks like LangChain, methods to mitigate hallucinations, and optimization via quantization. Learn to build autonomous agents for real-world use.\n\n### 🔑 Mind Map (Key Concepts)\n[📥 **Download High-Resolution Mind Map** (.jpg)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/building-llms-for-production/building-llms-for-production-mindmap.jpg)\n\u003cbr\u003e  \n\u003cdetails\u003e\n  \u003csummary\u003e\n    \u003cb\u003e\u003cem\u003e\u003ca\u003e🔍 Click here to unfold the full Mind Map (building-llms-for-production-mindmap.jpg)\u003c/a\u003e\u003c/em\u003e\n    \u003cbr\u003e (点击展开完整思维导图)\n    \u003c/b\u003e\n  \u003c/summary\u003e\n\n  ![Building LLMs for Production Mind Map](./summaries/building-llms-for-production/building-llms-for-production-mindmap.jpg)\n \n\u003c/details\u003e\n\n\n### 📑 Presentation Slides\n\u003e 💡 **Tip:** Press `Ctrl` + `Click` (or Command + Click) to open in a new tab.   \n[📥 View the \"Building LLMs for Production\" Slides (PDF)](./summaries/building-llms-for-production/building-llms-for-production-slides.pdf)   \n[📥 **Download PDF** (Direct Link)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/building-llms-for-production/building-llms-for-production-slides.pdf)\n  \n\n### 🛠️ Hands-on Lab \u0026 Examples  \n\n* **[DeepDive](https://github.com/Eric-LLMs/DeepDive):** Production-grade AI system for personalized learning — agentic tutoring, dual-track memory, hybrid RAG, personal/team cloud storage, and self-hosted LLM infrastructure. Highlights:\n  - **Config-driven RAG pipeline**: plug-and-play retrieval nodes *(query rewrite → vector + keyword recall → RRF fusion → cross-encoder rerank → parent-expand → relevance check)* that you add / remove / reorder / toggle live from the admin console — no code, no restart; chunking preview, golden-set Eval (Recall@k / Precision@k / MRR), degrade-never-break retrieval, vision-LLM PDF table transcription, a Redis query cache (keyed on query + config + corpus version, auto-invalidated on re-index), and user feedback logged to an eval dataset.\n  - **Agent kernel**: ReactLoopAgent step loop with long-term dual-track memory (PG tsvector + pgvector via RRF, recency-decay weighted) plus hierarchical history compaction (bounded context, flat token window), a skill catalog, hot-reloadable, dependency-injected plugins, and a read-only sandbox gating every tool call (READ / WRITE / NETWORK).\n  - **Reliability \u0026 human-in-the-loop**: hard timeout + exponential-backoff retry for transient LLM errors, per-turn cost budget, human-approval gate (Redis pub/sub, deny on timeout), plan mode + bounded subagents, shadow-git checkpoints, and a Docker bash sandbox (network-disabled, resource-capped).\n  - **Layered, cache-friendly prompts \u0026 tools**: byte-stable head (identity + one-line tool catalog) + per-step dynamic tail (memory / skills / injected) — warm prefix cache (lower latency \u0026 token cost), cache identity measurable; deferred tools — a searchable one-line catalog, matched tools mount as stable stubs, full schemas via the search result.\n  - **Cloud drive**: private **My Drive** + shared **workspaces** with owner / admin / editor / viewer roles, member management, and an append-only activity log; per-user object store with SHA-256 content-addressing (identical files deduplicate to one ref-counted blob), 8 MB resumable chunked upload, multi-level folders, trash \u0026 30-day retention, per-file ACLs and public links, and in-window previews for PDF / video / audio / Office documents.\n  - **Permissions \u0026 key management**: admin console for roles, users, LLM provider credentials + model catalog + routing weights, a per-user key-grant matrix (masked `sk-***`), SMTP, and stateless signed admin sessions.\n  - **Local-first \u0026 self-hosted**: Electron workbench works offline (file tree, multi-format viewer, video screenshots), big media is processed on the local client, and the whole stack (PostgreSQL/pgvector, Redis, TEI embedding, Kokoro TTS, LiteLLM gateway) runs via docker-compose — your data stays yours.\n\n* **[LLMs-Lab](https://github.com/Eric-LLMs/LLMs-Lab):** Research modules covering **Fine-Tuning**, **RAG optimization**, **LangChain**, **Prompt Engineering**, **Function-Calling**, **Agent**, etc. — each studied as a standalone module with multiple project implementations.\n\n[⬆️ Back to Top : Table of Contents](#top)  \n  \n---\n\u003cbr\u003e  \n\n\u003ca id=\"high-performance-private-ai-infrastructure\"\u003e\u003c/a\u003e   \n  \n# 📚 Building High-Performance, Private AI Infrastructure for the Enterprise\nCovers the full-stack AI infrastructure for the enterprise — AI chips, compute clusters, high-speed networking, distributed training, inference serving, cluster scheduling, and secure private deployment.\n\n### 🛠️ Hands-on Projects \u0026 Tools  \n\n**I. AI Infrastructure**\n\n* **[AIInfra — AI Infrastructure Reference](https://github.com/Infrasys-AI/AIInfra):** An open-source reference covering the full-stack AI infrastructure for LLMs — from AI chips, compute clusters, and high-speed networking, to distributed training, inference optimization, and deployment — a hands-on resource for building high-performance, private AI infrastructure for the enterprise.\n\n**II. Cluster Scheduling \u0026 Orchestration**\n\n* **[Volcano](https://github.com/volcano-sh/volcano):** A Kubernetes-native batch system with advanced GPU scheduling and job queueing — widely adopted for enterprise AI clusters.\n\n* **[KubeRay](https://github.com/ray-project/kuberay):** A Kubernetes operator for running Ray clusters, bridging distributed compute with cloud-native orchestration.\n\n**III. Training \u0026 Post-training**\n\n#### Pre-training\n\n* **[DeepSpeed](https://github.com/microsoft/DeepSpeed):** Microsoft's deep learning optimization library — ZeRO memory optimization, mixed precision, and system optimizations for training and fine-tuning models at massive scale.\n\n* **[Megatron-LM](https://github.com/NVIDIA/Megatron-LM):** NVIDIA's large-scale language model training framework — tensor, pipeline, and sequence parallelism, often paired with DeepSpeed for pre-training.\n\n#### Fine-tuning\n\n* **[Unsloth](https://github.com/unslothai/unsloth):** A fast, memory-efficient fine-tuning library — up to 2x faster and 70% less memory for LoRA/QLoRA fine-tuning of LLMs.\n\n* **[LLaMA-Factory](https://github.com/hiyouga/LLaMA-Factory):** A config-driven fine-tuning platform supporting LoRA, QLoRA, and full-parameter tuning across many open LLMs.\n\n* **[HuggingFace PEFT](https://github.com/huggingface/peft):** The standard parameter-efficient fine-tuning library — LoRA, QLoRA, and more — widely used for enterprise model customization.\n\n#### RL / Alignment\n\n* **[veRL](https://github.com/verl-project/verl):** ByteDance Seed's production-grade RL post-training framework — supports PPO, GRPO, DAPO, PRIME, and multi-turn tool-calling agents, with vLLM/SGLang rollout and FSDP/Megatron-LM training backends.\n\n* **[OpenRLHF](https://github.com/OpenRLHF/OpenRLHF):** A high-performance distributed RLHF framework built on Ray, vLLM, and DeepSpeed — supporting PPO, GRPO, REINFORCE++, and multi-turn agent training.\n\n* **[TRL](https://github.com/huggingface/trl):** HuggingFace's official library for RLHF and post-training — the broadest algorithm coverage (SFT, DPO, GRPO, PPO, RLOO) with first-party OpenEnv integration.\n\n* **[ART](https://github.com/OpenPipe/ART):** OpenPipe's Agent Reinforcement Trainer — a lightweight GRPO framework that adds RL training loops (inference → reward → LoRA) to any existing Python application.\n\n**IV. Inference Serving \u0026 Deployment**\n\n#### High-performance\n\n* **[vLLM](https://github.com/vllm-project/vllm):** A high-throughput, memory-efficient LLM serving engine (PagedAttention, continuous batching, prefix caching) — the de facto standard for high-performance private inference.\n\n* **[SGLang](https://github.com/sgl-project/sglang):** A fast, structured-generation runtime for LLM inference, complementing vLLM with radix attention and efficient prefix reuse.\n\n#### Lightweight\n\n* **[Ollama](https://github.com/ollama/ollama):** The simplest way to run LLMs locally — a lightweight, self-hosted private deployment option.\n\n* **[LocalAI](https://github.com/mudler/LocalAI):** A local, OpenAI-compatible, self-hosted inference server for private model deployment.\n\n**V. Distributed Computing**\n\n* **[Ray](https://github.com/ray-project/ray):** A unified distributed framework for AI training, inference, and serving at scale.\n\n**VI. Unified Gateway**\n\n* **[LiteLLM](https://github.com/BerriAI/litellm):** A unified LLM gateway with an OpenAI-compatible API — model routing, rate limits, budgets, and logging for enterprise private deployments.\n\n**VII. Security \u0026 Guardrails**\n\n* **[NeMo Guardrails](https://github.com/NVIDIA/NeMo-Guardrails):** NVIDIA's programmable guardrails framework for conversational AI — input, output, and retrieval rails.\n* **[LLM Guard](https://github.com/protectai/llm-guard):** Protect AI's input/output security library for detecting prompt injection and sanitizing LLM traffic.\n* **[PurpleLlama / Llama Guard](https://github.com/meta-llama/PurpleLlama):** Meta's Llama security toolkit — Llama Guard content-safety classifier and Prompt Guard injection detection.\n* **[Garak](https://github.com/NVIDIA/garak):** NVIDIA's LLM vulnerability scanner for automated red-teaming.\n* **[Microsoft Presidio](https://github.com/microsoft/presidio):** PII detection and data anonymization for compliance in enterprise AI deployments.\n\n[⬆️ Back to Top : Table of Contents](#top)  \n  \n---\n\u003cbr\u003e  \n\u003ca id=\"building-ai-agents\"\u003e\u003c/a\u003e  \n  \n# 📚 Building AI Agents\n  \n### 🔑 Mind Map (Key Concepts)\n[📥 **Download High-Resolution Mind Map** (.jpg)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/introduction-to-ai-agents/agents-architecture-operations-and-evolution-mindmap.jpg)\n  \u003cbr\u003e  \n\u003cdetails\u003e\n  \u003csummary\u003e\n    \u003cb\u003e\u003cem\u003e\u003ca\u003e🔍 Click here to unfold the full Mind Map (agents-architecture-operations-and-evolution-mindmap.jpg)\u003c/a\u003e\u003c/em\u003e \n    \u003cbr\u003e (点击展开完整思维导图)\n    \u003c/b\u003e\n  \u003c/summary\u003e  \n\n  ![Building AI Agents Mindmap](./summaries/introduction-to-ai-agents/agents-architecture-operations-and-evolution-mindmap.jpg)\n \n\u003c/details\u003e  \n  \n### 📑 Presentation Slides\n\u003e 💡 **Tip:** Press `Ctrl` + `Click` (or Command + Click) to open in a new tab.   \n[📥 View the \"Building AI Agents\" Slides (PDF)](./summaries/introduction-to-ai-agents/agents-architecture-operations-slides.pdf)   \n[📥 **Download PDF** (Direct Link)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/introduction-to-ai-agents/agents-architecture-operations-slides.pdf)  \n  \n\n### 🛠️ Hands-on: A Minimal ReAct Agent  \n👉 [**View the AI Agent Project in the LLMs-Lab repository on the Eric-LLMs GitHub profile.**](https://github.com/Eric-LLMs/LLMs-Lab/tree/main/Agent/Agent_Project)\n  \nTo bridge theory with practice, I developed a modular AI Agent project that implements autonomous reasoning and task execution:\n\n* **Architecture:** Utilizes a decoupled structure with dedicated directories for `Agent` logic, `Tools`, `Utils`, and `Prompts`.\n* **Reasoning Loop:** Features an `AutoGPT.py` implementation using **ReAct (Reasoning and Acting)** logic to handle complex, multi-step goal decomposition.\n* **Functional Tools:** Includes custom tools for deep data analysis (Excel processing via Pandas), automated communication via email, PDF-based QA interrogation (**FileQATool**), requirements-driven document generation (**WriterTool**), and dynamic script-based auditing of structured files using custom heuristics and thresholds (**PythonTool**).\n* **End-to-End Workflow:** Supports real-world scenarios, such as identifying underperforming suppliers from sales records and autonomously drafting/sending notifications.\n\n### 🧰 Key Open-Source Projects \u0026 References\n\nThe following open-source projects represent prominent examples of agentic AI engineering:\n\n| Project | Description | Key Strengths |\n| :--- | :--- | :--- |\n| **[Claude Code](https://github.com/anthropics/claude-code)** | Anthropic's official terminal-based agentic coding tool | Agentic coding, terminal-native, full codebase understanding, git workflows |\n| **[claurst](https://github.com/Kuberwastaken/claurst)** | Community-maintained reference implementation of Claude Code | Internal architecture study, reverse-engineering insights, codebase structure reference |\n| **[OpenAI Codex](https://github.com/openai/codex)** | OpenAI's open-source agentic coding CLI | Agentic coding, terminal-native, sandboxed execution, bash tool use |\n| **[Hermes-Agent](https://github.com/NousResearch/hermes-agent)** | Self-improving AI agent with built-in learning loop | Skill creation from experience, cross-session memory, multi-channel (CLI/Telegram/Discord/Slack) |\n| **[OpenClaw](https://github.com/openclaw/openclaw)** | Personal AI assistant, local-first, any OS/platform | Local-first Gateway, multi-channel messaging, voice support, session \u0026 tool management |\n| **[Pi](https://github.com/earendil-works/pi)** | Open-source AI agent toolkit: unified multi-provider LLM API, agent runtime, and an interactive coding agent CLI | Self-extensible coding agent, multi-provider API, terminal UI library, npm supply-chain hardening |\n| **[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)** | DeepSeek AI's open-source agent harness with a plugin-driven architecture (\"everything is a plugin\") | Plugin composability via Cordis, Web UI, monorepo, developer preview |\n\n\u003e These projects showcase diverse agent architectures — from developer-focused coding agents (Claude Code/OpenAI Codex/claurst) to general-purpose personal assistants (OpenClaw) and self-learning agents (Hermes-Agent). Studying their design decisions is valuable for building your own agent systems.\n  \n  \n[⬆️ Back to Top : Table of Contents](#top)  \n  \n---\n\u003cbr\u003e  \n\n\n\u003ca id=\"mastering-the-model-context-protocol\"\u003e\u003c/a\u003e   \n \n# 📚 Mastering the Model Context Protocol (MCP)  \nA deep dive into the Model Context Protocol (MCP) — the open standard that connects AI agents to tools and data sources. Covers the protocol architecture, official SDKs, and production server implementations.\n\n### 🔑 Mind Map (Key Concepts)\n[📥 **Download High-Resolution Mind Map** (.jpg)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/mastering-the-model-context-protocol/mastering-the-model-context-protocol-mindmap.jpg)\n\u003cbr\u003e  \n\u003cdetails\u003e\n  \u003csummary\u003e\n    \u003cb\u003e\u003cem\u003e\u003ca\u003e🔍 Click here to unfold the full Mind Map (mastering-the-model-context-protocol-mindmap.jpg)\u003c/a\u003e\u003c/em\u003e\n    \u003cbr\u003e (点击展开完整思维导图)\n    \u003c/b\u003e\n  \u003c/summary\u003e\n\n  ![Mastering the Model Context Protocol (MCP)](./summaries/mastering-the-model-context-protocol/mastering-the-model-context-protocol-mindmap.jpg)\n \n\u003c/details\u003e\n\n\n### 📑 Presentation Slides\n\u003e 💡 **Tip:** Press `Ctrl` + `Click` (or Command + Click) to open in a new tab.   \n[📥 View the \"Mastering the Model Context Protocol (MCP)\" Slides (PDF)](./summaries/mastering-the-model-context-protocol/mastering-the-model-context-protocol-slides.pdf)   \n[📥 **Download PDF** (Direct Link)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/mastering-the-model-context-protocol/mastering-the-model-context-protocol-slides.pdf)\n  \n\n### 🧰 Key Frameworks \u0026 Tools\n\nThe official Model Context Protocol SDKs and reference implementations for building and connecting MCP servers:\n* **[specification](https://github.com/modelcontextprotocol/modelcontextprotocol)**: The official protocol specification and schema — core primitives (tools, resources, prompts) and transports.\n* **[typescript-sdk](https://github.com/modelcontextprotocol/typescript-sdk)**: The official TypeScript SDK for building MCP servers and clients.\n* **[python-sdk](https://github.com/modelcontextprotocol/python-sdk)**: The official Python SDK for building MCP servers and clients.\n* **[servers — Official Reference Implementations](https://github.com/modelcontextprotocol/servers)**: The official collection of reference MCP servers, including filesystem, fetch, git, memory, and sequential thinking.\n* **[registry](https://github.com/modelcontextprotocol/registry)**: The official, community-driven MCP server registry — an \"app store\" for discovering standardized MCP servers.\n* **[FastMCP](https://github.com/jlowin/fastmcp)** *(community)*: The most popular high-level Python framework for building MCP servers — define tools and resources in a few lines of code.\n\n### 🔗 Related Protocols\n\nThe Model Context Protocol connects agents to tools and data. For agent-to-agent collaboration and client interfaces, see:\n* **[A2A (Agent2Agent Protocol)](https://github.com/a2aproject/A2A)**: The open standard for agent-to-agent communication — lets agents discover each other's capabilities and collaborate on tasks.\n* **[ACP (Agent Client Protocol)](https://github.com/coder/agent-client-protocol)**: A protocol for connecting agents to editors and frontends.\n\n### 🛠️ Hands-on Projects \u0026 Tools  \n\n👉 **[Explore Model Context Protocol (MCP) Projects on GitHub](https://github.com/Eric-LLMs/awesome-mcp-servers)** *A curated collection of industry-standard Model Context Protocol (MCP) server implementations.*\n\n[⬆️ Back to Top : Table of Contents](#top)  \n  \n---\n\u003cbr\u003e  \n\n\u003ca id=\"agent-memory-part-i\"\u003e\u003c/a\u003e   \n \n# 📚 Agent Memory Part I\nA survey of academic research on how agent memory is designed and categorized (forms, functions, dynamics).\n\n### 🔑 Mind Map (Key Concepts)\n[📥 **Download High-Resolution Mind Map** (.jpg)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/memory-in-the-age-of-ai-agents-survey/unforgettable_agents_architecting_ai_memory-mindmap.jpg)\n\u003cbr\u003e  \n\u003cdetails\u003e\n  \u003csummary\u003e\n    \u003cb\u003e\u003cem\u003e\u003ca\u003e🔍 Click here to unfold the full Mind Map (unforgettable_agents_architecting_ai_memory-mindmap.jpg)\u003c/a\u003e\u003c/em\u003e\n    \u003cbr\u003e (点击展开完整思维导图)\n    \u003c/b\u003e\n  \u003c/summary\u003e\n\n  ![Unforgettable Agents Architecting AI Memory](./summaries/memory-in-the-age-of-ai-agents-survey/unforgettable_agents_architecting_ai_memory-mindmap.jpg)\n \n\u003c/details\u003e\n\n\n### 📑 Presentation Slides   \n  \n#### A Blueprint for Memory in Agentic Intelligence\n\u003e 💡 **Tip:** Press `Ctrl` + `Click` (or Command + Click) to open in a new tab.   \n[📥 View the \"A Blueprint for Memory in Agentic Intelligence\" Slides (PDF)](./summaries/memory-in-the-age-of-ai-agents-survey/a-blueprint-for-memory-in-agentic-intelligence.pdf)   \n[📥 **Download PDF** (Direct Link)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/memory-in-the-age-of-ai-agents-survey/a-blueprint-for-memory-in-agentic-intelligence.pdf)\n  \n#### Unforgettable Agents Architecting AI Memory\n\u003e 💡 **Tip:** Press `Ctrl` + `Click` (or Command + Click) to open in a new tab.   \n[📥 View the \"Unforgettable Agents Architecting AI Memory\" Slides (PDF)](./summaries/memory-in-the-age-of-ai-agents-survey/unforgettable_agents_architecting_ai_memory.pdf)   \n[📥 **Download PDF** (Direct Link)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/memory-in-the-age-of-ai-agents-survey/unforgettable_agents_architecting_ai_memory.pdf)\n\n\n### 📑 Further Reading / Resources\n\nFor a comprehensive list of papers related to Agent Memory, we highly recommend checking out:  \n👉 [Agent-Memory-Paper-List](https://github.com/Shichun-Liu/Agent-Memory-Paper-List) by Shichun-Liu.\n\n\n[⬆️ Back to Top : Table of Contents](#top)  \n  \n---\n\u003cbr\u003e  \n\n  \n\u003ca id=\"building-memory-modules-for-agentic-ai-Systems\"\u003e\u003c/a\u003e   \n \n# 📚 Building Memory Modules for Agentic AI Systems\nA comprehensive guide on designing memory systems for AI Agents. This document synthesizes academic surveys with practical implementation strategies — covering the taxonomy of agent memory (forms, functions, dynamics), deep dives into Mem0, Letta (MemGPT), and LangMem, and enterprise-grade solutions using Amazon Bedrock AgentCore.  \n\n### 🔑 Mind Map (Key Concepts)\n[📥 **Download High-Resolution Mind Map** (mindmap.png)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/building-memory-for-agentic-ai-theory-frameworks-and-practice/building-memory-for-agentic-ai-theory-frameworks-and-practice-mindmap.png)\n\u003cbr\u003e  \n\u003cdetails\u003e\n  \u003csummary\u003e\n    \u003cb\u003e\u003cem\u003e\u003ca\u003e🔍 Click here to unfold the full Mind Map\u003c/a\u003e\u003c/em\u003e\n    \u003cbr\u003e (点击展开完整思维导图)\n    \u003c/b\u003e\n  \u003c/summary\u003e\n\n  ![memory solution in production](./summaries/building-memory-for-agentic-ai-theory-frameworks-and-practice/building-memory-for-agentic-ai-theory-frameworks-and-practice-mindmap.png)\n \n\u003c/details\u003e\n\n\n### 📑 Presentation Slides   \n  \n#### Building Memory for Agentic AI: Theory, Frameworks, and Practice\n\u003e 💡 **Tip:** Press `Ctrl` + `Click` (or Command + Click) to open in a new tab.   \n[📥 View Slides (PDF)](./summaries/building-memory-for-agentic-ai-theory-frameworks-and-practice/building-memory-for-agentic-ai-theory-frameworks-and-practice.pdf)   \n[📥 **Download PDF** (Direct Link)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/building-memory-for-agentic-ai-theory-frameworks-and-practice/building-memory-for-agentic-ai-theory-frameworks-and-practice.pdf)\n\n### 🧰 Key Frameworks \u0026 Code Samples\n\nThe following frameworks and repositories are discussed in this guide, representing the current state-of-the-art in Agentic Memory:  \n* **[Mem0](https://github.com/mem0ai/mem0)**: A dual-layer memory framework supporting working, factual, and semantic memory types for agent state persistence.\n* **[Letta (MemGPT)](https://github.com/letta-ai/letta)**: Manages infinite context by treating agents like an OS with virtual memory and recursive summarization.\n* **[LangMem](https://github.com/langchain-ai/langmem)**: A LangChain library that implements Semantic, Episodic, and Procedural memory integration for LangGraph agents.\n* **[Zep / Graphiti](https://github.com/getzep/graphiti)**: Zep's temporal knowledge-graph framework — builds a dynamic, time-aware memory graph for agent state with causal event support.\n* **[Amazon Bedrock Samples](https://github.com/aws-samples/amazon-bedrock-samples)**: A comprehensive collection of examples for using Amazon Bedrock, including various implementations of Agentic workflows and memory patterns.\n  \n[⬆️ Back to Top : Table of Contents](#top)  \n  \n---\n\u003cbr\u003e  \n\n\u003ca id=\"agent-eval\"\u003e\u003c/a\u003e   \n  \n# 📚 Agent Evaluation (Eval) Engineering\n\nEvaluating AI Agents requires a fundamental shift from simple output checks (\"vibe checks\") to analyzing multi-step trajectories, environment changes, and tool usage. This repository consolidates frameworks and engineering practices for moving from **intuition to instrumentation**.\n\n---\n\n### 🔑 Key Considerations\n\n| Consideration | Description |\n| :--- | :--- |\n| The Intuition Trap | Why manual \"vibe checks\" fail as complexity scales. |\n| The Harness | Building a standardized environment for agent execution composed of Inputs, Tasks, and Graders. |\n| Trajectory vs. Outcome | Evaluating the *journey* (reasoning logs, tool calls) rather than just the *destination* (final answer). |\n| Reliability Metrics — Pass@k | Can the agent succeed *at least once* in k tries? (Good for brainstorming). |\n| Reliability Metrics — Pass^k | Can the agent succeed *every single time* in k tries? (Critical for autonomous agents). |\n| Swiss Cheese Model | Layering defenses (Automated Evals → Human Review → Production Monitoring) to ensure reliability. |\n| LLM-as-a-Judge | Using LLMs to grade outputs — with known biases (position bias, self-preference) that need calibration. |\n| Task Benchmarks | Standardized suites (SWE-bench, GAIA, AgentBench, τ-bench, WebArena) for measuring real-world task success. |\n| Tool-Call Correctness | Verifying the right tool, right arguments, and right timing — beyond just the final answer. |\n| Process Supervision | Grading intermediate reasoning and tool-call steps, not only the outcome, to catch errors early. |\n| Adversarial Robustness | Stress-testing against prompt injection and goal hijacking. |\n| Agent Security | Testing permission boundaries, tool authorization, and data-handling safety — ensuring the agent cannot overstep access or leak sensitive data. |\n| Cost \u0026 Latency | Token efficiency, wall-clock time, and per-task budget — decisive for production agents. |\n| Long-Horizon Tasks | Sustained multi-step planning and memory over long-running tasks. |\n\n### 🔑 Mind Map (Framework Overview)\n[📥 **Download High-Resolution Mind Map** (mindmap.png)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/agent-evaluation/ai-agent-evaluation-framework.png)\n\u003cbr\u003e  \n\u003cdetails\u003e\n  \u003csummary\u003e\n    \u003cb\u003e\u003cem\u003e\u003ca\u003e🔍 Click here to unfold the full Mind Map\u003c/a\u003e\u003c/em\u003e\n    \u003cbr\u003e (点击展开完整思维导图)\n    \u003c/b\u003e\n  \u003c/summary\u003e\n\n  ![Agent Evaluation Framework](./summaries/agent-evaluation/ai-agent-evaluation-framework.png)\n \n\u003c/details\u003e\n\n---\n\n### 📑 Presentation Slides\nA comprehensive guide to evaluating AI agents, focusing on the engineering framework for testing — including the \"Clean Room\" methodology, reliability metrics (Pass@k), and the \"Harness\" architecture. It treats evaluation as a core development practice.\n\n\u003e 💡 **Tip:** Press `Ctrl` + `Click` (or Command + Click) to open in a new tab.   \n[📥 View Slides (PDF)](./summaries/agent-evaluation/agent-evaluation-engineering.pdf)   \n[📥 **Download PDF** (Direct Link)](https://raw.githubusercontent.com/Eric-LLMs/Awesome-AI-Engineering/main/summaries/agent-evaluation/agent-evaluation-engineering.pdf)\n\n---\n\n### 🧰 Key Tools, Frameworks \u0026 Strategies\n\n#### 1. The Tooling Stack (Ecosystem)\nImplementing a robust evaluation pipeline requires specific infrastructure. The following tools are referenced and utilized in this framework:\n\n| Tool | Category | Key Features |\n| :--- | :--- | :--- |\n| **[LangSmith](https://smith.langchain.com/)** | Tracing \u0026 Debugging | Full trajectory tracing, `runnableConfig` tagging for A/B testing, and dataset management. |\n| **[LangFuse](https://langfuse.com/)** | Observability | Open-source alternative for observability, prompt management, and lightweight evaluation. |\n| **[Arize Phoenix](https://github.com/Arize-ai/phoenix)** | Observability \u0026 Eval | Open-source LLM tracing, embedding analysis, and RAG/agent evaluation. |\n| **[W\u0026B Weave](https://github.com/wandb/weave)** | Tracing \u0026 Eval | Lightweight LLM instrumentation, eval harnesses, and dataset versioning. |\n| **[DeepEval](https://github.com/confident-ai/deepeval)** | Unit Testing | \"Pytest for LLMs\". Specific metrics for RAG (Hallucination, Answer Relevancy) and Agents. |\n| **[OpenAI Evals](https://github.com/openai/evals)** | Evaluation Framework | OpenAI's open-source framework for model-graded evals — YAML/JSON config-driven, custom eval classes, and dataset registries. |\n| **[Braintrust](https://github.com/braintrustdata/braintrust-sdk)** | Evaluation Platform | Dataset management, LLM-as-a-judge scoring, online evals, and A/B testing. |\n| **[OpenEvals](https://github.com/langchain-ai/openevals)** | Graders | A library of pre-built \"LLM-as-a-judge\" prompts (Conciseness, Correctness, Coherence) compatible with LangSmith. |\n| **[AgentOps](https://github.com/AgentOps-AI/agentops)** | Agent DevOps \u0026 Monitoring | Session replays, agent benchmarking, and cost \u0026 reliability tracking for autonomous agents. |\n| **[Promptfoo](https://github.com/promptfoo/promptfoo)** | Red Teaming \u0026 Regression | Declarative eval configs, LLM regression testing, and adversarial red-teaming. |\n\n#### 2. Architecture: Hybrid Agent (Fast vs. Slow)\nTo balance cost and performance, we implement a **Hybrid Agent Architecture**:\n* **Reactive Layer (System 1)**: Handles simple, direct queries (e.g., \"What is the stock price?\") with low latency.\n* **Deliberative Layer (System 2)**: Activated for complex planning or multi-step reasoning tasks.\n* **Coordination Layer**: A router that classifies intent and dispatches tasks.\n\nEach layer is evaluated with different metrics:\n* **Reactive Layer**: latency and single-step accuracy.\n* **Deliberative Layer**: task completion rate, multi-step planning correctness, and trajectory quality.\n* **Coordination Layer**: intent classification accuracy — misrouting is a common source of downstream failures.\n\n#### 3. Evaluation Strategy: The \"Clean Room\"\nTo prevent \"cheating\" through shared state, every evaluation trial runs in a fresh container/sandbox.\n* **Isolation**: Fresh container for every trial, plus state reset (environment, conversation history) and snapshot rollback to guarantee a clean slate.\n* **Mocking \u0026 Replay**: Simulate external APIs — or record-and-replay real responses — to control latency and produce deterministic, reproducible outputs.\n* **Determinism**: Fix seeds and use `temperature=0` so runs are repeatable and differences are attributable to code, not randomness.\n* **Cleanup \u0026 Anti-Leakage**: Aggressive state teardown (no shared history) and guard against goal leakage that could let the agent take shortcuts via shared state.\n\n[⬆️ Back to Top : Table of Contents](#top)\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/eric-llms%2Fawesome-ai-engineering/projects"}