{"id":52451223,"url":"https://github.com/doobidoo/mcp-memory-service","last_synced_at":"2026-08-28T03:15:12.402Z","repository":{"id":377292275,"uuid":"908539519","full_name":"doobidoo/mcp-memory-service","owner":"doobidoo","description":"Open-source persistent memory for AI agent pipelines (LangGraph, CrewAI, AutoGen) and Claude. REST API + knowledge graph + autonomous consolidation.","archived":false,"fork":false,"pushed_at":"2026-05-30T05:08:13.000Z","size":29529,"stargazers_count":1902,"open_issues_count":13,"forks_count":288,"subscribers_count":12,"default_branch":"main","last_synced_at":"2026-08-20T19:18:32.277Z","etag":null,"topics":["agent-memory","agentic-ai","ai-agents","autogen","claude","crewai","knowledge-graph","langgraph","long-term-memory","mcp","mcp-server","memory","model-context-protocol","multi-agent","open-source","rag","semantic-search","sqlite-vec","vector-database","vector-storage"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/doobidoo.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":".github/FUNDING.yml","license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":".github/CODEOWNERS","security":"SECURITY.md","support":null,"governance":null,"roadmap":"docs/ROADMAP.md","authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":"NOTICE","maintainers":null,"copyright":null,"agents":"AGENTS.md","claude":"CLAUDE.md","gemini":null,"cursor":null,"copilot":null,"dco":null,"cla":null,"disclosure":null},"funding":{"github":"doobidoo","ko_fi":"doobidoo","custom":["https://www.buymeacoffee.com/doobidoo","https://paypal.me/heinrichkrupp1"]}},"created_at":"2024-12-26T10:15:44.000Z","updated_at":"2026-08-20T09:49:57.000Z","dependencies_parsed_at":"2026-08-20T19:28:46.049Z","dependency_job_id":null,"html_url":"https://github.com/doobidoo/mcp-memory-service","commit_stats":null,"previous_names":["doobidoo/mcp-memory-service"],"tags_count":499,"template":false,"template_full_name":null,"purl":"pkg:github/doobidoo/mcp-memory-service","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/doobidoo%2Fmcp-memory-service","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/doobidoo%2Fmcp-memory-service/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/doobidoo%2Fmcp-memory-service/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/doobidoo%2Fmcp-memory-service/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/doobidoo","download_url":"https://codeload.github.com/doobidoo/mcp-memory-service/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/doobidoo%2Fmcp-memory-service/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36948614,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-08-28T02:00:06.244Z","response_time":114,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agent-memory","agentic-ai","ai-agents","autogen","claude","crewai","knowledge-graph","langgraph","long-term-memory","mcp","mcp-server","memory","model-context-protocol","multi-agent","open-source","rag","semantic-search","sqlite-vec","vector-database","vector-storage"],"created_at":"2026-08-19T15:00:22.702Z","updated_at":"2026-08-28T03:15:12.393Z","avatar_url":"https://github.com/doobidoo.png","language":"Python","funding_links":["https://github.com/sponsors/doobidoo","https://ko-fi.com/doobidoo","https://www.buymeacoffee.com/doobidoo","https://paypal.me/heinrichkrupp1"],"categories":["Memory Management","MCP服务器与插件","Python","MCP Servers \u0026 Protocol","Model Serving \u0026 Inference","🧠 Knowledge \u0026 Memory (62 servers)","MCP Ecosystem","Containerised MCP Servers"],"sub_categories":["Vector Databases \u0026 Retrieval Infrastructure","General-Purpose Machine Learning","Servers","Database \u0026 Storage"],"readme":"# mcp-memory-service\n\n## Persistent Shared Memory for AI Agent Pipelines\n\nOpen-source memory backend for AI agents — **REST API, MCP, OAuth, CLI, dashboard**. One self-hosted service, every transport.\nAgents store decisions, share causal knowledge graphs, and retrieve\ncontext in 5ms — without cloud lock-in or API costs.\n\n**Works with LangGraph · CrewAI · AutoGen · any HTTP client · Claude Desktop · OpenCode**\n\n---\n\n[![Website](https://img.shields.io/badge/Website-mcpmemory.services-00e5ff?logo=cloudflare\u0026logoColor=white)](https://mcpmemory.services)\n[![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)\n[![PyPI version](https://img.shields.io/pypi/v/mcp-memory-service?color=blue\u0026logo=pypi\u0026logoColor=white)](https://pypi.org/project/mcp-memory-service/)\n[![Python](https://img.shields.io/pypi/pyversions/mcp-memory-service?logo=python\u0026logoColor=white)](https://pypi.org/project/mcp-memory-service/)\n[![Codeberg stars](https://img.shields.io/gitea/stars/doobidoo/mcp-memory-service?gitea_url=https%3A%2F%2Fcodeberg.org\u0026logo=codeberg\u0026label=stars)](https://codeberg.org/doobidoo/mcp-memory-service)\n[![Works with LangGraph](https://img.shields.io/badge/Works%20with-LangGraph-green)](https://github.com/langchain-ai/langgraph)\n[![Works with CrewAI](https://img.shields.io/badge/Works%20with-CrewAI-orange)](https://crewai.com)\n[![Works with AutoGen](https://img.shields.io/badge/Works%20with-AutoGen-purple)](https://github.com/microsoft/autogen)\n[![Works with Claude](https://img.shields.io/badge/Works%20with-Claude-blue)](https://claude.ai)\n[![Works with Cursor](https://img.shields.io/badge/Works%20with-Cursor-orange)](https://cursor.sh)\n[![Remote MCP](https://img.shields.io/badge/MCP-Remote%20Support-blue?logo=anthropic)](docs/remote-mcp-setup.md)\n[![claude.ai Browser Compatible](https://img.shields.io/badge/claude.ai-Browser%20Compatible-orange?logo=anthropic)](docs/remote-mcp-setup.md)\n[![OAuth 2.0](https://img.shields.io/badge/Auth-OAuth%202.0%20%2B%20DCR-green)](docs/oauth-setup.md)\n\n---\n\n\u003cdiv align=\"center\"\u003e\n  \u003cvideo src=\"https://mcpmemory.services/assets/videos/knowledge-graph-3d.mp4\" poster=\"https://mcpmemory.services/assets/images/knowledge-graph-3d-poster.png\" width=\"820\" autoplay loop muted playsinline controls\u003e\n    \u003ca href=\"https://mcpmemory.services/\"\u003e\u003cimg src=\"docs/assets/images/knowledge-graph-3d.png\" alt=\"3D knowledge graph — memories as a glowing, interactive galaxy\" width=\"820\"\u003e\u003c/a\u003e\n  \u003c/video\u003e\n  \u003cp\u003e\u003cem\u003e▶ \u003ca href=\"https://mcpmemory.services/\"\u003eThe 3D knowledge graph in motion\u003c/a\u003e\u003c/em\u003e — every memory a glowing node, every relationship a curved edge. \u003csub\u003e(Video not playing? \u003ca href=\"https://mcpmemory.services/\"\u003eSee it live at mcpmemory.services\u003c/a\u003e.)\u003c/sub\u003e\u003c/p\u003e\n\u003c/div\u003e\n\n---\n\n## Why Agents Need This\n\nYour AI assistant forgets everything when you start a new chat. You spend 10 minutes re-explaining your architecture. **Again.** MCP Memory Service captures project context, architecture decisions, and code patterns automatically — new sessions start with everything already known.\n\n| Without mcp-memory-service | With mcp-memory-service |\n|---|---|\n| Each agent run starts from zero | Agents retrieve prior decisions in 5ms |\n| Memory is local to one graph/run | Memory is shared across all agents and runs |\n| You manage Redis + Pinecone + glue code | One self-hosted service, zero cloud cost |\n| No causal relationships between facts | Knowledge graph with typed edges (causes, fixes, contradicts) |\n| Context window limits create amnesia | Autonomous consolidation compresses old memories |\n\n**Key capabilities for agent pipelines:**\n- **Framework-agnostic REST API** — 76 endpoints, no MCP client library needed\n- **Knowledge graph** — agents share causal chains, not just facts\n- **`X-Agent-ID` header** — auto-tag memories by agent identity for scoped retrieval\n- **`conversation_id`** — bypass deduplication for incremental conversation storage\n- **SSE events** — real-time notifications when any agent stores or deletes a memory\n- **Embeddings run locally via ONNX** — memory never leaves your infrastructure\n\n---\n\n## 🚀 Get Started in 60 Seconds\n\n\u003e Not sure which setup fits your needs? See the **[Setup Guide](docs/setup-guide.md)** — a decision tree walks you to the right path in under a minute.\n\n**1. Install:**\n\n```bash\npip install mcp-memory-service\n```\n\n**2. Configure your AI client:**\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cstrong\u003eClaude Desktop\u003c/strong\u003e\u003c/summary\u003e\n\nAdd to your config file:\n- **macOS**: `~/Library/Application Support/Claude/claude_desktop_config.json`\n- **Windows**: `%APPDATA%\\Claude\\claude_desktop_config.json`\n- **Linux**: `~/.config/Claude/claude_desktop_config.json`\n\n```json\n{\n  \"mcpServers\": {\n    \"memory\": {\n      \"command\": \"memory\",\n      \"args\": [\"server\"]\n    }\n  }\n}\n```\n\nRestart Claude Desktop. Your AI now remembers everything across sessions.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eClaude Code\u003c/strong\u003e\u003c/summary\u003e\n\n```bash\nclaude mcp add memory -- memory server\n```\n\nRestart Claude Code. Memory tools will appear automatically.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eAgent pipelines (REST API — LangGraph, CrewAI, AutoGen, any HTTP client)\u003c/strong\u003e\u003c/summary\u003e\n\n```bash\nMCP_ALLOW_ANONYMOUS_ACCESS=true memory server --http\n# REST API running at http://localhost:8000\n```\n\n```python\nimport asyncio\nimport httpx\n\nBASE_URL = \"http://localhost:8000\"\n\n\nasync def main():\n    async with httpx.AsyncClient() as client:\n        # Store — auto-tag with X-Agent-ID header\n        await client.post(f\"{BASE_URL}/api/memories\", json={\n            \"content\": \"API rate limit is 100 req/min\",\n            \"tags\": [\"api\", \"limits\"],\n        }, headers={\"X-Agent-ID\": \"researcher\"})\n        # Stored with tags: [\"api\", \"limits\", \"agent:researcher\"]\n\n        # Search — scope to a specific agent\n        results = await client.post(f\"{BASE_URL}/api/memories/search\", json={\n            \"query\": \"API rate limits\",\n            \"tags\": [\"agent:researcher\"],\n        })\n        print(results.json()[\"memories\"])\n\n\nasyncio.run(main())\n```\n\n**Framework-specific guides:** [docs/agents/](docs/agents/)\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003eOpenCode\u003c/strong\u003e\u003c/summary\u003e\n\nStart the HTTP API:\n\n```bash\nMCP_ALLOW_ANONYMOUS_ACCESS=true memory server --http\n```\n\nInstall the local plugin:\n\n```bash\ngit clone https://codeberg.org/doobidoo/mcp-memory-service.git\ncd mcp-memory-service\nmkdir -p ~/.config/opencode/plugins\ncp opencode/memory-plugin.js ~/.config/opencode/plugins/\ncp opencode/memory-plugin.config.example.json ~/.config/opencode/memory-plugin.json\n```\n\nOpenCode automatically loads local plugins from `~/.config/opencode/plugins/` and `.opencode/plugins/`.\n\nOptional: register the `/memory` slash command in `~/.config/opencode/opencode.json` to query status, search, and health from inside the TUI:\n\n```json\n{\n  \"command\": {\n    \"memory\": {\n      \"description\": \"Show MCP Memory Service status. Usage: /memory, /memory search \u003cquery\u003e, /memory health\",\n      \"template\": \"\"\n    }\n  }\n}\n```\n\nSee [OpenCode integration guide](opencode/README.md) for configuration, project-local installs, slash command details, TUI toasts, and current limitations.\n\n\u003e The current OpenCode integration ships as repository files for the local plugin directory. If you installed only the PyPI package, clone the repository once to copy the plugin files.\n\u003e\n\u003e The plugin defaults to `http://127.0.0.1:8000`, but `memoryService.endpoint` and `OPENCODE_MEMORY_ENDPOINT` let you target any reachable HTTP deployment.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🌐 claude.ai (Browser — Remote MCP)\u003c/strong\u003e\u003c/summary\u003e\n\nUnlike desktop-only MCP servers, mcp-memory-service supports **Remote MCP**: persistent memory directly in your browser, on any device — no Claude Desktop required. Enterprise-ready (OAuth 2.0 + HTTPS + CORS), self-hosted or cloud-hosted.\n\n```bash\n# 1. Start server with Remote MCP\nMCP_STREAMABLE_HTTP_MODE=1 \\\nMCP_SSE_HOST=0.0.0.0 \\\nMCP_OAUTH_ENABLED=true \\\npython -m mcp_memory_service.server\n\n# 2. Expose publicly (Cloudflare Tunnel)\ncloudflared tunnel --url http://localhost:8765\n\n# 3. Add connector in claude.ai Settings → Connectors with the tunnel URL\n#    OAuth flow will handle authentication automatically\n```\n\n**Production Setup:** [Remote MCP Setup Guide](docs/remote-mcp-setup.md) (Let's Encrypt, nginx, Docker, firewall).\n**Step-by-Step Tutorial:** [Blog: 5-Minute claude.ai Setup](https://mcpmemory.services/blog/remote-mcp-tutorial.html) | [Wiki Guide](https://codeberg.org/doobidoo/mcp-memory-service/wiki/Claude-AI-Remote-MCP-Integration)\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🔧 Advanced: Custom Backends \u0026 Team Setup\u003c/strong\u003e\u003c/summary\u003e\n\nFor production deployments, team collaboration, or cloud sync:\n\n```bash\ngit clone https://codeberg.org/doobidoo/mcp-memory-service.git\ncd mcp-memory-service\npython scripts/installation/install.py\n```\n\nChoose from:\n- **SQLite** (local, fast, single-user)\n- **Cloudflare** (cloud, multi-device sync)\n- **Hybrid** (best of both: 5ms local + background cloud sync)\n- **Milvus** (dedicated vector DB — Milvus Lite file, self-hosted, or Zilliz Cloud)\n\n\u003e ℹ️ For long-lived services (MCP servers, web backends, notebook sessions), prefer Docker Milvus or Zilliz Cloud over Milvus Lite. See [docs/milvus-backend.md](docs/milvus-backend.md#which-uri-to-use) for why.\n\n\u003c/details\u003e\n\n---\n\n## ⚡ Works With Your Favorite AI Tools\n\n#### 🤖 Agent Frameworks (REST API)\n**LangGraph** · **CrewAI** · **AutoGen** · **Any HTTP Client** · **OpenClaw/Nanobot** · **Custom Pipelines**\n\n#### 🖥️ CLI \u0026 Terminal AI (MCP)\n**Claude Code** · **Gemini CLI** · **Gemini Code Assist** · **OpenCode** · **Codex CLI** · **Goose** · **Aider** · **GitHub Copilot CLI** · **Amp** · **Continue** · **Zed** · **Cody**\n\n#### 🎨 Desktop \u0026 IDE (MCP)\n**Claude Desktop** · **VS Code** · **Cursor** · **Windsurf** · **Kilo Code** · **Raycast** · **JetBrains** · **Replit** · **Sourcegraph** · **Qodo**\n\n#### 💬 Chat Interfaces (MCP)\n**ChatGPT** (Developer Mode) · **claude.ai** (Remote MCP via HTTPS)\n\n**Works seamlessly with any MCP-compatible client or HTTP client** - whether you're building agent pipelines, coding in the terminal, IDE, or browser.\n\n\u003e **💡 NEW**: ChatGPT now supports MCP! Enable Developer Mode to connect your memory service directly. [See setup guide →](docs/remote-mcp-setup.md)\n\n---\n\n## ✨ Features\n\n🧠 **Persistent Memory** – Context survives across sessions with semantic search\n🔍 **Smart Retrieval** – Finds relevant context automatically using AI embeddings\n⚡ **5ms Speed** – Instant context injection, no latency\n🔄 **Multi-Client** – Works across 25+ AI applications\n☁️ **Cloud Sync** – Optional Cloudflare backend for team collaboration\n🔒 **Privacy-First** – Local-first, you control your data\n📊 **Web Dashboard** – Visualize and manage memories at `http://localhost:8000`\n🧬 **Knowledge Graph** – Interactive D3.js visualization of memory relationships\n🏠 **Homelab Quality Scoring** – Point scoring at any OpenAI-compatible endpoint (Ollama, LiteLLM, vLLM)\n🔗 **Entity Extraction** – Auto-links @mentions, #tags, URLs, and file paths from memory content to a queryable entity graph\n💡 **Insight Cards** – Consolidation detects patterns, trends, and knowledge gaps across your memory corpus and surfaces them as structured insights\n🏷️ **Tag Match Filtering** – `tag_match=AND/OR` on `memory_search` for precise multi-tag queries\n\n### 🖥️ Dashboard Preview\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://codeberg.org/doobidoo/mcp-memory-service/wiki/raw/images/dashboard/mcp-memory-dashboard-v9.3.0-tour.gif\" alt=\"MCP Memory Dashboard Tour\" width=\"800\"/\u003e\n\u003c/p\u003e\n\n**8 Dashboard Tabs:** Dashboard • Search • Browse • Documents • Manage • Analytics • Quality • API Docs\n\n🎬 **[Watch the Web Dashboard Walkthrough on YouTube](https://youtu.be/W34r8VFoSdQ)** — semantic search, tag browser, document ingestion, analytics, quality scoring, and API docs in under 2 minutes.\n📖 See [Web Dashboard Guide](https://codeberg.org/doobidoo/mcp-memory-service/wiki/Web-Dashboard-Guide) for complete documentation.\n\n---\n\n## Real-World Deployments\n\n### Multi-Agent Cluster with Shared Memory\n\n\u003e *\"After I work with one of the cluster agents on something I want my local agent to know about, the cluster agent adds a special tag to the memory entry that my local agent recognizes as a message from a cluster agent. So they end up using it as a comms bridge — and it's pretty delightful.\"*\n\u003e — [@jeremykoerber](https://github.com/jeremykoerber) (originally GitHub issue #591)\n\nA 5-agent openclaw cluster uses mcp-memory-service as shared state **and** as an inter-agent messaging bus — without any custom protocol. Cluster agents tag memories with a sentinel like `msg:cluster`, and the local agent filters on that tag to receive cross-cluster signals. The memory service becomes the coordination layer with zero additional infrastructure.\n\n```python\n# Cluster agent stores a learning and flags it for the local agent\nawait client.post(f\"{BASE_URL}/api/memories\", json={\n    \"content\": \"Rate limit on provider X is 50 RPM — switch to provider Y after 40\",\n    \"tags\": [\"api\", \"limits\", \"msg:cluster\"],       # sentinel tag\n}, headers={\"X-Agent-ID\": \"cluster-agent-3\"})\n\n# Local agent polls for cluster messages\nresults = await client.post(f\"{BASE_URL}/api/memories/search\", json={\n    \"query\": \"messages from cluster\",\n    \"tags\": [\"msg:cluster\"],\n})\n```\n\nThis pattern — **tags as inter-agent signals** — emerges naturally from the tagging system and requires no additional infrastructure.\n\n### Self-Hosted Docker Stack with Cloudflare Tunnel\n\n\u003e *\"The quality of life that session-independent memory adds to AI workflows is immense. File-based memory demands constant discipline. Semantic recall from a live database doesn't. Storing data on my own hardware while making it remotely accessible across platforms turned out to be a feature I didn't know I needed.\"*\n\u003e — [@PL-Peter](https://github.com/PL-Peter) (originally GitHub discussion #602)\n\nA production-tested self-hosted deployment using Docker containers behind a Cloudflare tunnel, with [AuthMCP Gateway](https://github.com/loglux/authmcp-gateway) handling authentication:\n\n| Layer | Role |\n|-------|------|\n| **Cloudflare Tunnel** | Name-based routing, subnet-based access control, authentication before hitting self-hosted resources |\n| **AuthMCP Gateway** | Auth/aggregation with locally managed users, admin UI, per-user MCP server access control, bearer token auth |\n| **mcp-memory-service** | Two Docker containers sharing one SQLite backend — one for MCP, one for the web UI (document ingestion) |\n\n**Security best practices for this setup:**\n- Use Cloudflare ZeroTrust with subnet-based access control (e.g., allow Anthropic subnets + your own IPs)\n- Add **Client IP Address Filtering** to all Cloudflare API tokens (Dashboard → My Profile → API Tokens → Edit → Client IP Address Filtering) to limit abuse if a token leaks\n- If using IPv6, include your IPv6 /64 network in the allowlist (Python prefers IPv6 by default)\n- For long-running browser sessions, request the `offline_access` scope during authorization to receive a rotating `refresh_token` (lifetime via `MCP_OAUTH_REFRESH_TOKEN_EXPIRE_DAYS`, default 30 days). Without this scope, access tokens are the only credential — extend `MCP_OAUTH_ACCESS_TOKEN_EXPIRE_MINUTES` up to `1440` (24h) if you need longer single-shot sessions.\n- Consider an auth proxy like [AuthMCP](https://github.com/loglux/authmcp-gateway) or [mcp-auth-proxy](https://github.com/sigbit/mcp-auth-proxy) for robust session management\n\n### Fully-Offline Shared Memory Across Four Agents\n\n\u003e *\"mcp-memory-service has been the shared memory layer for all my coding agents since February — Claude Code, Claude Desktop, Codex CLI and OpenCode all talk to the same sqlite-vec DB over stdio on my Mac. ~5,900 memories and counting. Every session starts by pulling a bootstrap profile from memory and ends by committing a session summary, so any agent can pick up where another left off — work context, project state, even a 'mistakes I made before' log. It's the closest thing to persistent identity my agents have.\"*\n\u003e — Mingjian Shao (AI PM \u0026 AI consultant, via LinkedIn)\n\nA single local `sqlite-vec` database on a Mac acts as the shared brain for four different agents over stdio — no server, no cloud. Embeddings run fully offline via a local **Qwen3-Embedding-0.6B (1024-dim) on MPS**, with daily automated backups, scheduled consolidation, and the dashboard kept alive by a LaunchAgent.\n\n**Lesson worth stealing (offline embeddings):** when the custom embedding model fails to load, the service can silently fall back to the default MiniLM (384-dim) and subsequent writes fail with dimension mismatches. If you pin a non-default embedding model, also pin the model path and set the Hugging Face offline flags so a load failure surfaces loudly instead of degrading — then a dimension mismatch can't corrupt the store.\n\n---\n\n## Comparison with Alternatives\n\n### vs. Commercial Memory APIs\n\n| | Mem0 | Zep | DIY Redis+Pinecone | **mcp-memory-service** |\n|---|---|---|---|---|\n| License | Proprietary | Enterprise | — | **Apache 2.0** |\n| Cost | Per-call API | Enterprise | Infra costs | **$0** |\n| **🌐 claude.ai Browser** | ❌ Desktop only | ❌ Desktop only | ❌ | **✅ Remote MCP** |\n| **OAuth 2.0 + DCR** | ❓ Unknown | ❓ Unknown | ❌ | **✅ Enterprise-ready** |\n| **Streamable HTTP** | ❌ | ❌ | ❌ | **✅ (SSE also supported)** |\n| Framework integration | SDK | SDK | Manual | **REST API (any HTTP client)** |\n| Knowledge graph | No | Limited | No | **Yes (typed edges)** |\n| Auto consolidation | No | No | No | **Yes (decay + compression)** |\n| On-premise embeddings | No | No | Manual | **Yes (ONNX, local)** |\n| Privacy | Cloud | Cloud | Partial | **100% local** |\n| Hybrid search | No | Yes | Manual | **Yes (BM25 + vector)** |\n| MCP protocol | No | No | No | **Yes** |\n| REST API | Yes | Yes | Manual | **Yes (76 endpoints)** |\n\n### vs. MCP-Native Alternatives\n\n[MemPalace](https://github.com/MemPalace/mempalace) is an MCP-native alternative that went viral in April 2026 with strong LongMemEval claims. A [community code review (Issue #27)](https://github.com/MemPalace/mempalace/issues/27) subsequently showed that the headline numbers reflect the underlying vector store rather than the advertised Palace architecture, and the maintainers acknowledged most points. We keep the comparison here for transparency, but readers should interpret the scores with that context in mind.\n\n| | **MemPalace** | **mcp-memory-service** |\n|---|---|---|\n| LongMemEval R@5 (raw ChromaDB, zero LLM) | 96.6%¹ | 86.0% (session) / 80.4% (turn) |\n| LongMemEval R@5 (with reranking) | 100%² | — |\n| Storage granularity | Session-level | **Turn-level + session-level** |\n| Team / multi-device sync | ❌ Local only | **✅ Cloudflare sync** |\n| REST API / Web dashboard | ❌ | **✅** |\n| OAuth 2.1 + multi-user | ❌ | **✅** |\n| Knowledge graph | ❌ | **✅ (typed edges)** |\n| Auto consolidation | ❌ | **✅ (decay + compression)** |\n| Compatible AI tools | Claude-focused | **25+ tools** |\n| License | MIT | **Apache 2.0** |\n\n**Why the benchmark gap?** MemPalace stores whole sessions as single units — LongMemEval's \"which session contains the answer?\" question is answered structurally by that granularity. mcp-memory-service defaults to turn-level storage for fine-grained retrieval; using `memory_store_session` brings our score to **86.0% R@5**. And per Issue #27, the 96.6% headline measures a raw ChromaDB baseline with the Palace architecture inactive — an apples-to-apples architectural comparison is not possible with the published numbers.\n\n\u003e ¹ Measured in MemPalace \"raw mode\" (plain text in ChromaDB with default embeddings). Per [Issue #27](https://github.com/MemPalace/mempalace/issues/27), the Palace structural features are bypassed in this configuration.\n\u003e\n\u003e ² 100% result uses optional LLM reranking (~500 API calls) on a partially tuned test set. Clean held-out score (as reported by the maintainers): **98.4% R@5**.\n\n---\n\n## 📊 Retrieval Benchmarks\n\nThree benchmarks measure retrieval quality (all-MiniLM-L6-v2, 384d embeddings, zero LLM API calls):\n\n**LongMemEval** ([500 questions](https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned), ~45–62 distractor sessions per question):\n\n| Question Type | R@5 | R@10 | NDCG@10 | MRR |\n|---------------|-----|------|---------|-----|\n| **Overall** | **80.4%** | **90.4%** | **82.2%** | **89.1%** |\n| single-session-assistant | 100.0% | 100.0% | 99.3% | 99.1% |\n| knowledge-update | 84.6% | 96.8% | 86.2% | 95.5% |\n| single-session-user | 91.4% | 92.9% | 86.0% | 83.8% |\n| temporal-reasoning | 72.0% | 84.1% | 75.1% | 85.7% |\n| multi-session | 70.7% | 86.0% | 77.6% | 89.4% |\n\n**DevBench** (practical developer workflow queries):\n\n| Category | Recall@5 | MRR |\n|----------|----------|-----|\n| **Overall** | **91.1%** | **0.861** |\n| exact | 100% | 1.000 |\n| semantic | 80.0% | 0.700 |\n| cross-type | 90.0% | 0.867 |\n\n**LoCoMo** ([ACL 2024](https://github.com/snap-research/locomo) long-term conversational memory):\n\n| Category | Recall@5 | MRR |\n|----------|----------|-----|\n| **Overall** | **49.7%** | **0.414** |\n| multi-hop | 72.0% | 0.600 |\n| temporal | 33.5% | 0.274 |\n\nRun benchmarks: `python scripts/benchmarks/benchmark_longmemeval.py`, `python scripts/benchmarks/benchmark_devbench.py`, `python scripts/benchmarks/benchmark_locomo.py`\n\n---\n\n## 🛠️ Configuration Highlights\n\nFull reference: **[Configuration Guide](docs/mastery/configuration-guide.md)**\n\n### Server Lifecycle (CLI)\n\n```bash\nmemory launch                  # Start HTTP server in background (127.0.0.1:8000)\nmemory launch --port 8192      # Custom port\nmemory info                    # Status and health\nmemory logs --lines 50         # Recent logs\nmemory stop                    # Stop server\n```\n\nThese commands are optimized for fast startup and avoid loading heavy ML dependencies unless needed.\n\n\u003e ⚠️ **Security Note**: By default, the server binds to `127.0.0.1` (localhost only). `--host 0.0.0.0` / `MCP_HTTP_HOST=0.0.0.0` exposes the API to your network — do this only in trusted environments with proper authentication and firewall rules. For untrusted networks, use TLS termination (reverse proxy with HTTPS) or VPN overlays.\n\n### Embedding Model Selection\n\nThe default model (`all-MiniLM-L6-v2`) works well for **English-only** content. If you store memories in other languages, switch to a multilingual model:\n\n| Model | Languages | Dimensions | Use case |\n|-------|-----------|-----------|----------|\n| `all-MiniLM-L6-v2` (default) | English only | 384 | Fastest, English-only deployments |\n| `paraphrase-multilingual-MiniLM-L12-v2` | 50+ languages | 384 | Mixed-language or non-English content |\n\n```bash\nexport MCP_EMBEDDING_MODEL=paraphrase-multilingual-MiniLM-L12-v2\n```\n\n\u003e ⚠️ **Switching models requires re-embedding existing memories** (cross-language cosine drops from ~0.95 to ~0.10 otherwise): stop the service, run `python scripts/maintenance/regenerate_embeddings.py` with the new model env var, restart.\n\n### Quality Scoring with Your Local LLM\n\n**Homelab / self-hosted quality scoring** (v10.45.0+): set `MCP_QUALITY_AI_PROVIDER=openai-compatible` to score memories with your local LLM instead of ONNX or a cloud API:\n\n```bash\nMCP_QUALITY_AI_PROVIDER=openai-compatible\nMCP_QUALITY_AI_BASE_URL=http://localhost:11434/v1   # Ollama\nMCP_QUALITY_AI_MODEL=qwen2.5:7b-instruct\n# MCP_QUALITY_AI_API_KEY=ollama                     # optional\n```\n\nRecommended models: `qwen2.5:7b-instruct` (Ollama), `mlx-community/Qwen2.5-7B-Instruct-4bit` (MLX), or any instruct model via LiteLLM proxy. On endpoint failure, scoring falls back to implicit signals automatically.\n\n**Local quality scoring in a container.** The standard and `:slim` images ship `onnxruntime` but not the exported ONNX models, and not the `torch`/`transformers` needed to export them — so `MCP_QUALITY_AI_PROVIDER=local` needs the models supplied from outside. There is no published `:quality-cpu` tag; it was retired rather than rebuilt per release, because the ONNX models are version-independent and rebuilding them on every patch was waste. Three supported paths:\n\n1. **Export once, mount the directory** (recommended). Run `scripts/quality/export_deberta_onnx.py` on any machine with `torch`/`transformers`, then mount the result and point `MCP_QUALITY_ONNX_MODEL_DIR` at it.\n2. **Build the image yourself** — `tools/docker/Dockerfile.quality-cpu` stays in the tree and does the export at build time.\n3. **Use an endpoint you already run** — `MCP_QUALITY_AI_PROVIDER=openai-compatible` against Ollama, vLLM, or a LiteLLM proxy, as configured above. No models to manage.\n\nRecipes for all three, including a verified non-root read-only-rootfs Kubernetes setup: [`tools/docker/README.md`](tools/docker/README.md).\n\n---\n\n## 🌐 SHODH Ecosystem Compatibility\n\nMCP Memory Service is **fully compatible** with the [SHODH Unified Memory API Specification v1.0.0](https://github.com/varun29ankuS/shodh-memory/blob/main/specs/openapi.yaml): all SHODH implementations share the same memory schema (emotional metadata, episodic memory, source tracking, quality scoring), so memories export/import across implementations with full fidelity.\n\n| Implementation | Backend | Embeddings | Use Case |\n|----------------|---------|------------|----------|\n| **[shodh-memory](https://github.com/varun29ankuS/shodh-memory)** | RocksDB | MiniLM-L6-v2 (ONNX) | Reference implementation |\n| **shodh-cloudflare** | Cloudflare Workers + Vectorize | Workers AI (bge-small) | Edge deployment, multi-device sync |\n| **mcp-memory-service** (this) | SQLite-vec / Hybrid | MiniLM-L6-v2 (ONNX) | Desktop AI assistants (MCP) |\n\n---\n\n## 📰 In the Media\n\n**[Agents Overdrawn at the Memory Bank](https://www.linkedin.com/pulse/humans-loop-deep-dive-agents-overdrawn-rbrcc/)** — Heavybit's *Humans in the Loop* deep dive talks to maintainer Heinrich Krupp about agent amnesia, why persistent memory is the missing infrastructure layer for agentic systems, and how mcp-memory-service closes the gap with local vector storage, ONNX embeddings, and typed knowledge graphs.\n\n\u003e \"Your project has inspired me in many ways. In my view, it's the best implementation of MCP memory I've found so far.\"\n\u003e — **Michał Zubkowicz**\n\n**[AI Tinkerers Zürich Talk](https://zurich.aitinkerers.org/talks/rsvp_DuHc2MBp1uo)** ([full video](https://drive.google.com/file/d/1NJDLetJ5CE7Jiek9OtyXCrt-zm768YB6/view)) — Maintainer Heinrich Krupp presents mcp-memory-service to the AI Tinkerers Zürich meetup, covering persistent memory architecture, semantic search, and multi-agent memory sharing.\n\n---\n\n## Latest Release: **v11.9.0** (August 27, 2026)\n\n**MINOR: transformers 5.x closes two high-severity advisories with no 4.x fix, plus a quality-system bug the new ml-extras CI job caught on day one**\n\n**What's New:**\n- **deps: move to transformers 5.x, closing two high-severity advisories** (#305). GHSA-29pf-2h5f-8g72 and GHSA-fgcw-684q-jj6r are both fixed only in 5.x. The reason to upgrade if you install the `[ml]` or `[nli]` extras.\n- **fix(quality): stop a disabled quality system from loading the ONNX ranker** (#316). `MCP_MEMORY_QUALITY_ENABLED=false` still triggered a full `torch.onnx.export` of DeBERTa on the first call — found when it blew pytest's 120s timeout in CI. Closes #315.\n- **fix(quality): let huggingface_hub resolve its own cache** (#307). `onnx_ranker` built cache paths by hand instead of letting the library resolve them, missing `HF_HOME`/`HF_HUB_CACHE` and sometimes picking an incomplete download. Closes #304.\n- **fix(deps): correct the setuptools bound that keeps milvus-lite importable** (#300).\n- **deps: put onnxscript in the `[ml]` extra so the quality export can actually run** (#310).\n- **ci: run the suite once with the ml and nli extras installed** (#303) — transformers and sentence-transformers had never been imported in CI before this.\n\n**Previous Releases** (v11 series — full history for all earlier versions in [CHANGELOG.md](CHANGELOG.md)):\n- **v11.8.5** - PATCH: two Docker fixes reproduced against the published images, plus a timezone-boundary bug in timeframe deletion, external contributor (#295, #297, #237, #298) (August 25, 2026)\n- **v11.8.4** - PATCH: hybrid deployment was burning through Cloudflare's D1 free-tier read allowance, plus three smaller correctness fixes (#289, #290, #287) (August 25, 2026)\n- **v11.8.3** - PATCH: reachable MCP transport advisory (CVE-2026-52869), HTTPS silently downgraded to HTTP on every restart, API key written to the access log (#277, #279, #285) (August 24, 2026)\n- **v11.8.2** - PATCH: OAuth `client_credentials` bypassed the owner API key (GHSA-5p27-64mv-pr73, CVSS 9.1) (August 23, 2026)\n- **v11.8.1** - PATCH: eight fixes on top of v11.8.0, six from timkjr — OAuth issuer validation and Docker HTTPS behaviour (#239, #231) (August 22, 2026)\n- **v11.8.0** - MINOR: the knowledge-graph layer actually works now — entity extraction was discarding every memory tag, and two features were gated on a storage attribute nothing ever set (#218, #219) (August 9, 2026)\n- **v11.7.0** - MINOR: three TLS certificate-verification bypasses gated behind explicit opt-in, a committed credential removed (#198, #210, #197/#200) (August 5, 2026)\n- **v11.6.1** - PATCH: harvest classifier provider chain fix (#180), Claude Code plugin manifest at 1.0.2 (#195) (August 3, 2026)\n- **v11.6.0** - MINOR: migration no longer drops the knowledge graph and derived beliefs when re-embedding (#189), locale-aware NER/NLI via YAML plugins (#54), Docker images ship the maintenance and migration scripts (#188) (August 2, 2026)\n- **v11.5.5** - PATCH: standard Docker image ships tokenizers so the ONNX backend actually loads (#162, #163, #164) (July 24, 2026)\n- **v11.5.4** - PATCH: web dashboard GitHub references replaced with Codeberg (#158, #159, @sunnyagain) (July 22, 2026)\n- **v11.5.3** - PATCH: Claude Code hooks config resolution under Marketplace install + graph orphan-prune `has_entity` fix + belief-derivation noise filter (#155, #156, #150, #151, #121, #152, @filhocf, @tecnobrat) (July 22, 2026)\n- **v11.5.2** - PATCH: sqlite_vec `delete_memory` proxy fix + hash-embedding fallback guard + embedding-dimension mismatch guard (#140, #135, #143, @jonatanbellido, @nxxxsooo) (July 15, 2026)\n- **v11.5.1** - PATCH: multi-store migration dimension safety + embedding-backend verification in `memory status` (#134, #136, @nxxxsooo) (July 15, 2026)\n- **v11.5.0** - MINOR: conditional temporal decay + functional belief derivation + consolidation clustering fix + bootstrap belief injection (#123, #124, #126, #127, @filhocf) (July 10, 2026)\n- **v11.4.0** - MINOR: memory merge action + pluggable domain NER extractors + mcpmemory.services landing page (#100, #54, @filhocf) (July 4, 2026)\n- **v11.3.3** - PATCH: fix(cli): memory CLI commands respect MCP_HTTPS_ENABLED (fixes silent failures when TLS is enabled) (July 1, 2026)\n- **v11.3.2** - PATCH: declare numpy\u003e=1.24.0 as core dependency (fixes uvx bare install crash, closes #98) (June 30, 2026)\n- **v11.3.1** - PATCH: claude-hooks noise reduction - auto-capture moved to Stop event and gated on substantive content (June 22, 2026)\n- **v11.3.0** - MINOR: Interactive 3D knowledge graph visualization (Orrery-inspired), node cap 100→1000/max 500→10000 (June 21, 2026)\n- **v11.2.0** - MINOR: OAuth security hardening (#91), sqlite-vec rowid collision fix (#90), composite graph scoring (#55/#77, @filhocf), OpenCode XDG state dir fix (#84) (June 20, 2026)\n- **v11.1.0** - MINOR: two-phase query API aggregation + maintenance script hardening (PR #78, @filhocf) (June 18, 2026)\n- **v11.0.0** - MAJOR: legacy tool-name alias removal + optional ML dependencies / ONNX-first fallback (PR #72, #49, #71) (June 13, 2026)\n\n**Full version history**: [CHANGELOG.md](CHANGELOG.md) | [Older versions (v10.36.3 and earlier)](docs/archive/CHANGELOG-HISTORIC.md) | [All Releases](https://codeberg.org/doobidoo/mcp-memory-service/releases)\n\n---\n\n## 📚 Documentation \u0026 Resources\n\n- **[Agent Integration Guides](docs/agents/)** – LangGraph, CrewAI, AutoGen, HTTP generic\n- **[OpenCode Integration](opencode/README.md)** – Local plugin for memory retrieval and context injection\n- **[Remote MCP Setup (claude.ai)](docs/remote-mcp-setup.md)** – Browser integration via HTTPS + OAuth\n- **[Setup Guide](docs/setup-guide.md)** – Decision tree + step-by-step paths for all use cases\n- **[Configuration Guide](docs/mastery/configuration-guide.md)** – Backend options and customization\n- **[Architecture Overview](docs/architecture.md)** – How it works under the hood\n- **[Team Setup Guide](docs/setup-guide.md#path-4-full-stack)** – OAuth and cloud collaboration\n- **[Token-Efficient Retrieval](docs/guides/token-efficient-retrieval.md)** – Bounding search responses (`limit`, `max_response_chars`) and the `memory_explore` → `memory_detail` knowledge map\n- **[Knowledge Graph Dashboard](docs/features/knowledge-graph-dashboard.md)** – Interactive graph visualization guide\n- **[Memory Type Ontology](docs/memory-ontology.md)** – Built-in taxonomy and `MCP_CUSTOM_MEMORY_TYPES` env var\n- **[Migration Guide](docs/MIGRATION.md)** – Upgrading between major versions (v9+ migrations run automatically on restart)\n- **[Troubleshooting](docs/troubleshooting/)** – Common issues and solutions\n- **[Technical Video Demo (2 min)](https://www.youtube.com/watch?v=veJME5qVu-A)** – Performance, architecture, AI/ML intelligence\n- **[API Reference](https://codeberg.org/doobidoo/mcp-memory-service/wiki)** – Programmatic usage\n- **[Wiki](https://codeberg.org/doobidoo/mcp-memory-service/wiki)** – Complete documentation\n- [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/doobidoo/mcp-memory-service) – AI-powered documentation assistant\n- **[MCP Starter Kit](https://kruppster57.gumroad.com/l/glbhd)** – Build your own MCP server using the patterns from this project\n\n---\n\n## 🤝 Contributing\n\nWe welcome contributions! See [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines.\n\nWho authors this project, who holds copyright, and what every change passes before\nit reaches `main`: [AUTHORSHIP.md](AUTHORSHIP.md).\n\n**Quick Development Setup:**\n```bash\ngit clone https://codeberg.org/doobidoo/mcp-memory-service.git\ncd mcp-memory-service\npip install -e .  # Editable install\npytest tests/      # Run test suite\n```\n\n---\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdoobidoo%2Fmcp-memory-service","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdoobidoo%2Fmcp-memory-service","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdoobidoo%2Fmcp-memory-service/lists"}