{"id":51826205,"url":"https://github.com/tcconnally/perseus-amd-act-ii","last_synced_at":"2026-07-22T10:04:00.199Z","repository":{"id":370370202,"uuid":"1293626065","full_name":"tcconnally/perseus-amd-act-ii","owner":"tcconnally","description":"Perseus Vault x AMD Instinct — encrypted, local-first agent memory kept off the GPU. AMD Developer Hackathon Act II (Unicorn Track).","archived":false,"fork":false,"pushed_at":"2026-07-09T01:58:45.000Z","size":3911,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-09T02:16:04.912Z","etag":null,"topics":["agent-memory","ai-agents","amd","fireworks-ai","hackathon","mcp","mi300x","perseus-vault","rocm","rust"],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/tcconnally.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","dco":null,"cla":null}},"created_at":"2026-07-08T13:20:20.000Z","updated_at":"2026-07-09T01:58:49.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/tcconnally/perseus-amd-act-ii","commit_stats":null,"previous_names":["tcconnally/perseus-amd-act-ii"],"tags_count":null,"template":false,"template_full_name":null,"purl":"pkg:github/tcconnally/perseus-amd-act-ii","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tcconnally%2Fperseus-amd-act-ii","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tcconnally%2Fperseus-amd-act-ii/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tcconnally%2Fperseus-amd-act-ii/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tcconnally%2Fperseus-amd-act-ii/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/tcconnally","download_url":"https://codeload.github.com/tcconnally/perseus-amd-act-ii/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tcconnally%2Fperseus-amd-act-ii/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35757404,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-22T02:00:06.236Z","response_time":124,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agent-memory","ai-agents","amd","fireworks-ai","hackathon","mcp","mi300x","perseus-vault","rocm","rust"],"created_at":"2026-07-22T10:03:59.242Z","updated_at":"2026-07-22T10:04:00.192Z","avatar_url":"https://github.com/tcconnally.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"assets/thumbnail.png\" alt=\"Perseus Vault x AMD Instinct - encrypted memory for AI agents, kept off the GPU\" width=\"100%\"\u003e\n\u003c/div\u003e\n\n# Perseus Vault × AMD Instinct\n\n\u003e **Encrypted, local-first, persistent memory for AI agents — kept off the GPU so\n\u003e every byte of MI300X HBM serves tokens.**\n\n[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](./LICENSE)\n[![Track: Unicorn](https://img.shields.io/badge/AMD%20Act%20II-Unicorn%20Track-ed1b2f.svg)]()\n[![ROCm](https://img.shields.io/badge/base-rocm%2Fdev--ubuntu--22.04-ed1b2f.svg)]()\n[![Reproducible](https://img.shields.io/badge/benchmarks-reproducible-3fb950.svg)]()\n[![Live demo](https://img.shields.io/badge/live%20demo-amd--demo.perseus.observer-3fb950.svg)](https://amd-demo.perseus.observer)\n\n**AMD Developer Hackathon: Act II — Unicorn (Open) Track.** Built on\n[Perseus Vault](https://github.com/Perseus-Computing-LLC/perseus-vault) — a\n**shipping** MIT-licensed memory engine (v2.20, 35 releases, 56 MCP tools, live on PyPI\nand the MCP registry), not a weekend build. lablab project:\n[lablab.ai/…/perseus](https://lablab.ai/ai-hackathons/amd-developer-hackathon-act-ii/perseus).\n\n\u003e ### ▶ Judges, start here — try the live demo: **[amd-demo.perseus.observer](https://amd-demo.perseus.observer)**\n\u003e Teach the agent a fact, open a brand-new session, and recall it — then watch an\n\u003e **open-weight LLM (gpt-oss-120b) answer live via the Fireworks AI API** using *only*\n\u003e what it recalled. Recall + footprint run on the **host CPU (0 bytes of GPU HBM)**;\n\u003e the demo's cross-vendor economics table shows the projection model (MI300X and\n\u003e 2×H100 are now both measured — see the banner below). Run a decay tick too. No login,\n\u003e per-visitor sandbox, daily budget cap on inference.\n\u003e\n\u003e \u003csub\u003eFireworks AI is this hackathon's designated inference partner — the\n\u003e [organizers describe the credits](https://lablab.ai/ai-hackathons/amd-developer-hackathon-act-ii)\n\u003e as access to *\"models hosted on AMD-hardware\"*, and Fireworks is\n\u003e [partnering with AMD](https://fireworks.ai/blog/fireworks-amd-ai-infrastructure-partnership)\n\u003e to serve on Instinct accelerators. Like any serving API it does not attest which\n\u003e accelerator handles a given request, so we do not claim one ourselves. The memory\n\u003e layer never touches a GPU regardless.\u003c/sub\u003e\n\n\u003e ### ⚠️ Honesty banner (please read)\n\u003e We rented a real **AMD Instinct MI300X** and measured the claims everything rests\n\u003e on, serving **Qwen2.5-72B** on vLLM/ROCm (host: AMD EPYC 9474F). Measured (medians):\n\u003e one card holds **15.3 concurrent 72B agents** at **$0.143/agent-hour** (validating\n\u003e our $0.133 projection); sustained **658 output tok/s** (peak 1,088) = **$0.92/1M\n\u003e output tokens** at retail rental — untuned bf16, no FP8/AITER, a floor not a ceiling;\n\u003e and the load-bearing result: with the MI300X **saturated serving the 72B**, recall on\n\u003e the host CPU moved **±0.6% (median of 6 runs, 18.7 → 18.8 ms p50)**. The memory layer\n\u003e steals ~zero inference cycles, proven under real load\n\u003e ([BENCHMARKS §3a](docs/BENCHMARKS.md#3a-measured-on-a-real-mi300x--data_source-measured)).\n\u003e We then rented **2× H100 SXM** and measured the cross-vendor claim too: best-case\n\u003e 5.0 concurrent agents at $1.68/agent-hr vs the MI300X's 15.3 at $0.143 —\n\u003e **11.7× measured-vs-measured, vs 2× H100** (a single H100 cannot load the model at all).\n\u003e We also rented and measured **8× A100 40GB** (57.9 agents, $0.275/agent-hr — but that's\n\u003e 8 cards to the MI300X's 1, so MI300X still wins 1.9× on $/agent-hr and 2.1× per card).\n\u003e And **2× A100 80GB** (RunPod, eager@0.97 — the exact 2× H100 config): 6.37 agents at\n\u003e $0.436/agent-hr, MI300X 3.0× cheaper. **Every cross-vendor row is now measured — no\n\u003e projection remains.** Every number is tagged `data_source`:\n\u003e **`measured`** (timed live, reproducible now), **`published-spec`** (vendor\n\u003e datasheet / cloud price list, cited below), or **`projection`** (derived from\n\u003e published-spec inputs with stated assumptions). **No projected number is presented\n\u003e as measured.** The demo and benchmark scripts print this warning on every run.\n\n---\n\n## Problem\n\nAn AI agent is only as smart as what it remembers — but its \"memory\" is usually just\nthe LLM context window, which dies when the session ends. The common fix is to bolt on\na vector database: embed everything, store the vectors, do nearest-neighbour search at\nrecall. That buys persistence at a steep price:\n\n- a **second system of record** (Postgres/pgvector, Pinecone, a Docker sidecar) that\n  drifts out of sync with the agent's state;\n- an **embedding model on the hot path** for every write and query — latency and GPU\n  cycles spent before the *actual* model runs;\n- **HBM pressure** — if the index or embedder shares the accelerator, it eats the very\n  memory you wanted for weights and KV cache.\n\nOn an AMD Instinct MI300X, that last point is the whole game. Its **192 GB of HBM3** is\nthe scarce resource. Every gigabyte spent storing what the agent *knows* is a gigabyte\nnot serving tokens.\n\n**Two markets feel this hardest.** Teams **bleeding tokens** re-feeding the same context\ninto every session; and **regulated orgs** — finance, defense, healthcare, the ones\nalready banning cloud AI tools — that legally *cannot* send agent memory to a cloud API\nat all. A vector-DB-in-the-cloud serves neither well. An encrypted, local-first,\noff-the-GPU memory layer serves both.\n\n## Solution\n\n**Perseus Vault** is a single Rust binary that gives an agent durable memory over the\nModel Context Protocol (MCP). Its recall path is **SQLite + FTS5 hybrid search** (BM25\nlexical ranking blended with a recency/decay prior) — **no embedding model, no external\nvector database, no GPU.**\n\nThat's the design, not a limitation. The memory layer runs on the **host CPU** beside\nthe accelerator, so:\n\n- **100% of the MI300X's 192 GB HBM3 stays available for weights + KV cache.**\n- Recall adds **zero GPU work** and never competes with inference for the accelerator.\n- **One GPU backs many concurrent agents** — each with its own AES-256-GCM-encrypted\n  memory file (~85 MB RAM + ~45 MB disk per 100K memories, **measured**), because those\n  files live in host RAM/disk, not HBM.\n\n## Architecture\n\n```\n                     AMD Instinct MI300X (192 GB HBM3)\n                  +------------------------------------+\n user turn ─────► | LLM weights + KV cache (inference) |\n     ▲            | via Fireworks AI / vLLM / ROCm     |\n     │            +------------------------------------+\n     │ recall THEN infer     ▲ grounding │ tokens\n     │                    ┌───────────────┐\n     └────────────────────┤  Agent loop   │\n                          └───────────────┘\n                             ▲ remember() / recall() / decay()  (MCP, CPU only)\n                    ┌───────────────────────────────┐\n                    │ Perseus Vault (Rust binary)    │\n                    │ SQLite+FTS5 · AES-256-GCM      │  ── 0 bytes HBM\n                    │ one portable .db file / agent  │\n                    └───────────────────────────────┘\n                          host CPU + RAM + disk\n```\n\nFull write-up: **[docs/ARCHITECTURE.md](docs/ARCHITECTURE.md)**.\n\n## Benchmarks\n\nFull tables, sources, and reproduction steps: **[docs/BENCHMARKS.md](docs/BENCHMARKS.md)**.\nReproduce §1–§2 with `python3 src/benchmark.py`.\n\n### Recall latency stays in low milliseconds as the store grows 100× — `measured`\nReference implementation (`src/benchmark.py`, AMD-CPU laptop, Python 3.14):\n\n| Entries | Recall p50 (ms) | Recall p99 (ms) | Insert ops/s |\n|---|---|---|---|\n| 1,000   | 0.20  | 0.39  | 72,917 |\n| 10,000  | 1.14  | 1.44  | 72,217 |\n| 100,000 | 11.87 | 15.67 | 68,276 |\n\nShipping engine (Perseus Vault v2.20.0, **measured**, AMD CPU — see\n[PERF.md](https://github.com/Perseus-Computing-LLC/perseus-vault/blob/main/PERF.md)):\nFTS5 recall **15.7 ms p50 / 23.2 ms p99** @100K, **hybrid recall 79.7 ms p50**\n(**3.7× faster since v2.19** — [#511](https://github.com/Perseus-Computing-LLC/perseus-vault/blob/main/PERF.md)),\nbi-temporal `as_of` **~0.1 ms flat**; bulk insert **98,732 entities/s**.\n\n### Footprint stays tiny — `measured`\n\n| Entries | DB file (MB) | RSS (MB, shipping engine) |\n|---|---|---|\n| 1,000   | 0.31  | — |\n| 10,000  | 2.61  | — |\n| 100,000 | 25.95 | ~85 |\n\n### One accelerator serves N agents — `measured` on MI300X **and** 2×H100\n\n**Measured on both sides (2026-07-09, same model, same vLLM 0.19.1, n=3 medians —\n[BENCHMARKS §3a–3b](docs/BENCHMARKS.md#3a-measured-on-a-real-mi300x--data_source-measured)):**\n\n| Serving Qwen2.5-72B bf16 | 1× MI300X | 2× A100 80GB | 2× H100 SXM | 8× A100 40GB |\n|---|---|---|---|---|\n| GPUs | **1** | 2 | 2 | 8 |\n| Config | standard | eager @0.97 | eager @0.97 (best case) | standard |\n| Concurrent 8K agents | **15.3** | **6.37** | **5.0** | **57.9** (7.2/card) |\n| **GPU $/agent-hour** | **$0.143** | **$0.436** | **$1.68** | **$0.275** |\n| $ / 1M output tokens | **$0.92** | $1.04 | $3.42 | $1.92 |\n\n*All four rows measured (Qwen2.5-72B bf16, vLLM 0.19.1). The 2× A100 80 GB and 2× H100 rows use the identical eager@0.97 config for a fair 160 GB comparison; a single H100 can't load the model at all. The 8× A100 is 320 GB (overprovisioned) — strong per-agent cost but eight cards vs one.*\n\n**11.7× lower $/agent-hour and 3.7× lower $/token vs 2× H100 — measured, not projected.**\nAt the *identical* configuration the H100 pair serves **zero** 8K requests (KV exhausted);\nits 5.0 figure is the best case that boots. (H100 does win per-stream decode latency, 39 vs\n83 ms TPOT — stated plainly.) We also **rented and measured 8× A100 40GB** (Lambda):\n57.9 concurrent agents at **$0.275/agent-hr** — but that is eight cards to the MI300X's one,\nso per-card the MI300X still leads (15.3 vs 7.2 agents/card) and wins **1.9×** on\n$/agent-hour. The **2× A100 80GB** config is now **measured** too (RunPod, eager@0.97 —\nthe same config as the 2× H100): 6.37 agents at $0.436/agent-hr, so the MI300X is 3.0×\ncheaper (the old ~$0.47 projection was validated). Perseus Vault memory runs on the CPU\n(~$0.0004/agent-hr ≈ 0.3% of the agent's cost) and uses **0 bytes of HBM**. Reproduce:\n`python3 src/economics.py` (projection model) + [BENCHMARKS §3a–3b](docs/BENCHMARKS.md)\n(measured runs, exact commands).\n\n## Quick start\n\n```bash\ngit clone https://github.com/tcconnally/perseus-amd-act-ii.git\ncd perseus-amd-act-ii\n\n# 1) Run the agent end-to-end (stdlib only, no GPU, no network needed):\npython3 src/agent_memory_demo.py\n\n# 2) Reproduce the measured benchmark tables + economics:\npython3 src/benchmark.py            # add --quick to skip the 100K row\n\n# 3) (optional) Real inference via the Fireworks AI API (open-weight model):\ncp .env.example .env                # then set FIREWORKS_API_KEY\npython3 src/agent_memory_demo.py\n\n# 4) (optional) Run against the real Perseus Vault Rust binary:\ncurl -sSf https://raw.githubusercontent.com/Perseus-Computing-LLC/perseus-vault/main/scripts/install.sh | sh\nPERSEUS_VAULT_BIN=~/.local/bin/perseus-vault python3 src/agent_memory_demo.py\n```\n\n### Docker (ROCm base — GPU-ready)\n\n```bash\ndocker build -t perseus-amd-act-ii .          # FROM rocm/dev-ubuntu-22.04:6.2\ndocker run --rm perseus-amd-act-ii            # runs the demo\ndocker run --rm perseus-amd-act-ii python3 src/benchmark.py --quick\n\n# On an AMD GPU host, expose the accelerator:\ndocker run --rm --device=/dev/kfd --device=/dev/dri --group-add video \\\n  -e FIREWORKS_API_KEY=... perseus-amd-act-ii\n```\n\n## Published-Spec Estimates\n\nGPU rows above are **not measured**. They come from vendor datasheets and 2026 cloud\nprice lists:\n\n- **AMD Instinct MI300X** — 192 GB HBM3, 5.325 TB/s bandwidth, 1,307.4 TFLOPS FP16\n  (peak), 750 W TDP.\n  [AMD MI300X datasheet (PDF)](https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/data-sheets/amd-instinct-mi300x-data-sheet.pdf) ·\n  [product page](https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html) ·\n  ROCm software: [rocm.docs.amd.com](https://rocm.docs.amd.com/).\n- **NVIDIA H100 SXM** — 80 GB HBM3, 3.35 TB/s, ~989 TFLOPS FP16, 700 W.\n- **NVIDIA A100 80GB SXM** — 80 GB HBM2e, 2.039 TB/s, 312 TFLOPS FP16, 400 W.\n- **Cloud pricing (2026, per-GPU-hour):** MI300X median ~$2.72 (from ~$1.99); H100\n  ~$3.93; A100 80GB ~$1.80. Sources: cloud-GPU price trackers\n  ([getdeploying](https://getdeploying.com/reference/cloud-gpu),\n  [thundercompute](https://www.thundercompute.com),\n  [gpucost.org](https://gpucost.org)), July 2026. Spot prices vary — but the headline\n  cross-vendor ratio no longer depends on tracker prices: it is **measured** at the\n  rates we actually paid ($2.19 MI300X, $8.38 2×H100 → 11.7×,\n  [BENCHMARKS §3b](docs/BENCHMARKS.md)). Even at a findable ~$2.85/hr per H100\n  ($5.70 for the pair), the measured agent ceilings give $1.14 vs $0.143 → **8.0×**.\n- **Model assumption:** Llama-3.1-70B, FP16 weights ~141 GB; KV cache per 8K-token\n  sequence ~2.5 GB (80 layers, 8 GQA KV heads, head_dim 128, fp16). Derivation lives in\n  `src/economics.py`.\n\n## What We Measured on Real AMD Hardware — and What's Next\n\nWe rented real MI300X time (twice) and measured the load-bearing claims; the rest stay\non the list (details in\n[docs/BENCHMARKS.md §4](docs/BENCHMARKS.md#4-what-we-measured-on-real-amd-hardware--and-whats-next)):\n\n1. **✅ Done — recall on the host EPYC CPU while the MI300X is busy.** +0.6% under a\n   synthetic 100%-utilization matmul (97.4 TFLOPS FP16,\n   [§1](docs/BENCHMARKS.md#1-recall-throughput--latency--data_source-measured)) *and*\n   **±0.6% (median, 6 runs) under a real vLLM serving load of Qwen2.5-72B**\n   ([§3a](docs/BENCHMARKS.md#3a-measured-on-a-real-mi300x--data_source-measured)).\n   The CPU memory layer steals ~no accelerator cycles.\n2. **✅ Done — true concurrent-agent ceiling: 15.3** measured from vLLM's KV-cache\n   budget serving a 72B on one MI300X (vs the ~20 idealized projection) — §3a.\n3. End-to-end agent-turn latency (CPU recall + MI300X generation) vs a vector-DB baseline.\n4. **✅ Done — measured $/agent-hour: $0.143** ($2.19/GPU-hr ÷ 15.3 agents) — §3a.\n   *Still open:* peak serving throughput → measured $/1M tokens (our current throughput\n   data is a single-process floor, so we don't headline it).\n5. A ROCm/HIP prototype offloading Perseus Vault's dense re-rank to an idle GPU slice —\n   an open question we'd answer with data, not claims.\n\n## Bonus building block: Gemma on AMD — the whole agent on one chip\n\nThe hackathon's partner challenge asks for the best **AMD-hosted Gemma** project. On\nFireworks, Gemma is *on-demand* — you deploy it yourself, and even the cheapest option\n(Gemma 4 E4B) bills **~$7/hour while idle**. Rather than pay to keep a model warm, we did\nthe more on-thesis thing and **self-hosted Gemma on AMD silicon for $0**:\n[`src/gemma_on_amd.py`](src/gemma_on_amd.py) runs the same recall→infer architecture with\n**Gemma 3 (4B-it, Q4_K_M GGUF) served locally by llama.cpp on an AMD CPU**, right beside\nthe Perseus Vault memory layer — the fleet story scaled down to a single chip, with no\nidle-billing meter running.\n\nMeasured on an **AMD Ryzen 7 9800X3D** (`measured`, reproduce with the script): recall\n**0.21 ms** + Gemma generation **~13 tok/s wall-clock** — no GPU, no cloud, no API key.\nOne architecture across the AMD lineup: **Gemma on a Ryzen/EPYC host** for single-agent\nboxes, **a 70B-class model on MI300X** for fleets — and the memory layer never moves.\n\n```bash\n# 1) llama.cpp:  winget install ggml.llamacpp   (or: brew install llama.cpp)\n# 2) an open Gemma GGUF:\ncurl -LO https://huggingface.co/ggml-org/gemma-3-4b-it-GGUF/resolve/main/gemma-3-4b-it-Q4_K_M.gguf\n# 3) serve + run:\nllama-server -m gemma-3-4b-it-Q4_K_M.gguf --port 8081 --ctx-size 8192 \u0026\npython3 src/gemma_on_amd.py       # prints your CPU + measured numbers\n```\n\n## What's in this repo\n\n| Path | |\n|---|---|\n| [`src/agent_memory_demo.py`](src/agent_memory_demo.py) | End-to-end stateful agent (learn → recall → infer → decay). |\n| [`src/gemma_on_amd.py`](src/gemma_on_amd.py) | Bonus: Gemma 3 + Perseus Vault on one AMD CPU (partner challenge). |\n| [`src/perseus_vault_store.py`](src/perseus_vault_store.py) | Memory interface: CPU reference store + real-binary bridge. |\n| [`src/benchmark.py`](src/benchmark.py) | Measured throughput/footprint tables + economics. |\n| [`src/economics.py`](src/economics.py) | The \"one MI300X serves N agents\" model. |\n| [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md) | Design + the off-the-GPU thesis. |\n| [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | All tables, sources, reproduction. |\n| [`docs/SUBMISSION.md`](docs/SUBMISSION.md) | Every lablab form field, pre-filled. |\n| [`Dockerfile`](Dockerfile) | ROCm-based, GPU-ready container. |\n\n## Not a weekend hack — a shipping product\n\nMost hackathon entries are born this week. Perseus Vault is a real product we brought\n*to* AMD — which is why the memory layer here is production-grade, not a prototype:\n\n- **Mature:** v2.20, **35 releases**, single ~8 MB Rust binary, **56 MCP tools**, AES-256-GCM.\n- **Distributed everywhere agents live:** published to **PyPI** as five framework adapters —\n  [LangChain](https://pypi.org/project/langchain-perseus-vault/),\n  [CrewAI](https://pypi.org/project/crewai-perseus-vault/),\n  [PydanticAI](https://pypi.org/project/pydantic-ai-perseus-vault/),\n  [Haystack](https://pypi.org/project/perseus-vault-haystack/),\n  [Google ADK](https://pypi.org/project/adk-perseus-vault-memory/) — and listed in the\n  **MCP registry**, **Smithery**, and **Glama** (`server.json` / `smithery.yaml` / `glama.json`).\n- **Running in production today.** (The live demo above runs this repo's CPU reference\n  implementation of the same recall path — see `webdemo/` — not the Rust binary.)\n\n### Why this matters to AMD\n\nAgent memory is a real, growing market (Mem0, Letta, Zep). Today those stateful-agent\nworkloads default to NVIDIA. Perseus Vault removes the reason they'd have to: by keeping\nmemory **off the accelerator**, it makes the MI300X's 192 GB HBM3 the cheapest place to\nrun a *fleet* of durable agents (**measured $0.143/agent-hr on a real MI300X — 11.7×\nunder a measured 2×H100 baseline**,\nsee [benchmarks](docs/BENCHMARKS.md)). And because the memory is local-first, **air-gap\nmode loses nothing** — the regulated buyers who most need on-prem get the *full* product,\nnot a degraded one (many stateful-agent tools quietly disable their best features\noffline; ours don't). In one line: **Perseus Vault turns AMD Instinct into the economical\nhome for the agent economy** — an adoption wedge for Instinct, not just another memory tool.\n\n## About\n\nBuilt by **Perseus Computing LLC** (Wyoming). Perseus Vault is MIT-licensed and\nproduction-deployed. This submission is original and MIT-compliant.\n\n## License\n\n[MIT](./LICENSE) © 2026 Perseus Computing LLC.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftcconnally%2Fperseus-amd-act-ii","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ftcconnally%2Fperseus-amd-act-ii","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftcconnally%2Fperseus-amd-act-ii/lists"}