{"id":49294986,"url":"https://github.com/cdeust/automatised-pipeline","last_synced_at":"2026-04-26T03:01:41.177Z","repository":{"id":351873864,"uuid":"1208941390","full_name":"cdeust/automatised-pipeline","owner":"cdeust","description":"Codebase intelligence as an MCP server — tree-sitter AST → LadybugDB graph → Louvain communities → hybrid BM25 + TF-IDF + RRF search. 23 tools · 10 stages · 220 tests · Rust · Clean Architecture. The read-only intelligence layer between finding and PRD.","archived":false,"fork":false,"pushed_at":"2026-04-25T14:17:12.000Z","size":1798,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-04-25T15:22:24.607Z","etag":null,"topics":["anthropic","bm25","claude","claude-code","claude-code-plugin","clean-architecture","code-intelligence","codebase-analysis","community-detection","cypher","graph-database","hybrid-search","louvain-algorithm","mcp-server","model-context-protocol","property-graph","rust","static-analysis","tree-sitter","zetetic"],"latest_commit_sha":null,"homepage":"https://ai-architect.tools","language":"Rust","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/cdeust.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-04-12T23:58:19.000Z","updated_at":"2026-04-25T14:17:14.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/cdeust/automatised-pipeline","commit_stats":null,"previous_names":["cdeust/automatised-pipeline"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/cdeust/automatised-pipeline","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cdeust%2Fautomatised-pipeline","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cdeust%2Fautomatised-pipeline/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cdeust%2Fautomatised-pipeline/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cdeust%2Fautomatised-pipeline/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/cdeust","download_url":"https://codeload.github.com/cdeust/automatised-pipeline/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cdeust%2Fautomatised-pipeline/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":32284333,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-25T18:29:39.964Z","status":"online","status_checked_at":"2026-04-26T02:00:05.962Z","response_time":129,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["anthropic","bm25","claude","claude-code","claude-code-plugin","clean-architecture","code-intelligence","codebase-analysis","community-detection","cypher","graph-database","hybrid-search","louvain-algorithm","mcp-server","model-context-protocol","property-graph","rust","static-analysis","tree-sitter","zetetic"],"created_at":"2026-04-26T03:01:37.849Z","updated_at":"2026-04-26T03:01:41.164Z","avatar_url":"https://github.com/cdeust.png","language":"Rust","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n  \u003cimg src=\"assets/banner.svg\" alt=\"automatised-pipeline — codebase intelligence as an MCP server\" width=\"100%\"/\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"LICENSE\"\u003e\u003cimg src=\"https://img.shields.io/badge/License-MIT-blue.svg\" alt=\"MIT License\"\u003e\u003c/a\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Rust-1.94+-dea584.svg\" alt=\"Rust 1.94+\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Tools-23-orange\" alt=\"23 MCP tools\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Tests-220_passing-brightgreen\" alt=\"220 tests\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Languages-Rust_·_Python_·_TypeScript-blueviolet\" alt=\"Languages\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Stages-0_through_9-8A2BE2\" alt=\"Stages\"\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"#what-an-agent-can-ask-it\"\u003eWhat An Agent Can Ask\u003c/a\u003e · \u003ca href=\"#getting-started\"\u003eGetting Started\u003c/a\u003e · \u003ca href=\"#the-pipeline\"\u003ePipeline\u003c/a\u003e · \u003ca href=\"#23-mcp-tools\"\u003eTools\u003c/a\u003e · \u003ca href=\"#architecture\"\u003eArchitecture\u003c/a\u003e · \u003ca href=\"#the-zetetic-standard\"\u003eZetetic Standard\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cstrong\u003eCompanion projects:\u003c/strong\u003e\u003cbr\u003e\n  \u003ca href=\"https://github.com/cdeust/Cortex\"\u003eCortex\u003c/a\u003e — persistent memory that consolidates and reconsolidates across sessions\u003cbr\u003e\n  \u003ca href=\"https://github.com/cdeust/zetetic-team-subagents\"\u003ezetetic-team-subagents\u003c/a\u003e — 97 genius reasoning agents + 18 team specialists\u003cbr\u003e\n  \u003ca href=\"https://github.com/cdeust/prd-spec-generator\"\u003eprd-spec-generator\u003c/a\u003e — TypeScript PRD generator that consumes our graph intelligence\n\u003c/p\u003e\n\n---\n\nEvery AI coding assistant hits the same wall: you ask it to change `handle_tool_call`, and it either hallucinates a function that was renamed last week, edits something in the wrong community of the codebase, or silently breaks a call chain three modules away. Agents operate on strings; codebases have structure. The gap is where bugs live.\n\n**automatised-pipeline** is a Rust MCP server that indexes any Rust / Python / TypeScript codebase into a LadybugDB property graph, resolves imports and call chains across files, detects functional communities via Leiden-class community detection, traces execution flows from entry points, builds a hybrid BM25 + sparse TF-IDF + RRF search index, and exposes all of it to AI agents through 23 MCP tools.\n\nIt is the **codebase intelligence layer** that sits between a finding (\"this bug exists\") and a PRD (\"here is the fix, here is what it affects, here is what it must never break\"). It is **read-only intelligence** — it never writes code, opens PRs, or runs CI. It tells the system what is true about the code so the next stage can reason without guessing.\n\n**One pipeline stage = one MCP tool. 10 stages. 23 tools. 12,000+ lines of Rust. 220 tests. Zero warnings. Every constant sourced.**\n\n---\n\n## What an agent can ask it\n\n```\nanalyze_codebase(path: \"/path/to/project\", output_dir: \"/tmp/run\")\n  → index + resolve + cluster + build search index in one call\n  → 430 nodes, 400 edges, 216 communities, 35 processes on our own codebase\n\nsearch_codebase(graph_path, query: \"process incoming tool requests\")\n  → hybrid ranked results: BM25 lexical + sparse TF-IDF semantic + RRF fusion\n  → returns: handle_tool_call (score 0.021), dispatch_request (0.020), ...\n\nget_context(graph_path, qualified_name: \"src/main.rs::handle_tool_call\")\n  → 360° view: community membership, process participation,\n    incoming calls, outgoing calls, types used, types that use it\n  → did-you-mean suggestions when the symbol isn't found exactly\n\nget_impact(graph_path, qualified_name)\n  → blast radius: every process that transits this symbol, every community it touches\n  → the answer to \"what breaks if I change this?\"\n\ndetect_changes(graph_path, diff_text OR base_ref+head_ref)\n  → git diff → affected symbols → impacted communities → touched processes\n  → risk score for the change\n\nvalidate_prd_against_graph(prd_path, graph_path)\n  → does the PRD reference real symbols? (symbol hallucination check)\n  → does \"scoped to X\" match the actual community count?\n  → does \"doesn't affect main\" hold against the call graph?\n\ncheck_security_gates(graph_path, changed_symbols)\n  → auth-critical community touch · unsafe symbol · public API change ·\n    unresolved imports · test coverage gap\n\nverify_semantic_diff(before_graph_path, after_graph_path)\n  → what nodes/edges appeared, what disappeared, what dangles,\n    new cycles via Tarjan SCC, regression score with verdict\n```\n\n---\n\n## Getting started\n\n### Prerequisites\n\n- Rust 1.94+ (`rustup install stable`)\n- CMake (LadybugDB builds its C++ core from source — ~5 minutes first build, cached after)\n\n### Clone + build\n\n```bash\ngit clone https://github.com/cdeust/automatised-pipeline.git\ncd automatised-pipeline\ncargo build --release\n# First build: ~5 minutes (compiles LadybugDB C++ core)\n# Subsequent builds: \u003c1 second incremental\n```\n\n### Register the MCP server\n\nThe repo ships a `.mcp.json` that Claude Code picks up automatically when you open the directory:\n\n```json\n{\n  \"mcpServers\": {\n    \"ai-architect\": {\n      \"command\": \"cargo\",\n      \"args\": [\"run\", \"--quiet\", \"--release\", \"--manifest-path\", \"Cargo.toml\"]\n    }\n  }\n}\n```\n\nOr register globally:\n\n```bash\nclaude mcp add ai-architect -- /absolute/path/to/target/release/ai-architect-mcp\n```\n\n### First run\n\n```bash\n# Run the binary directly to verify the handshake\n./target/release/ai-architect-mcp\n\n# Or exercise it via stdio JSON-RPC:\nprintf '%s\\n' \\\n  '{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"initialize\",\"params\":{}}' \\\n  '{\"jsonrpc\":\"2.0\",\"id\":2,\"method\":\"tools/list\"}' \\\n  '{\"jsonrpc\":\"2.0\",\"id\":3,\"method\":\"tools/call\",\"params\":{\"name\":\"health_check\",\"arguments\":{}}}' \\\n  | ./target/release/ai-architect-mcp\n```\n\n---\n\n## The pipeline\n\nEvery stage is a tool. Stages build on each other but are independently callable. The pipeline is serial in logical order but MCP calls are stateless — you can re-run stages 3a-3d on a fresh codebase without re-running stages 1-2.\n\n| # | Tool(s) | What it does |\n|---|---|---|\n| **0** | `health_check` | Handshake + protocol + tool count |\n| **1** | `extract_finding`, `refine_finding` | Deterministic finding extraction + orchestrator-aware prompt refinement |\n| **2** | `start_verification`, `append_clarification`, `finalize_verification`, `abort_verification` | Human-gated clarification loop with SHA-256 transcript digest, atomic single-file session state |\n| **3a** | `index_codebase`, `query_graph`, `get_symbol` | tree-sitter AST → LadybugDB graph (16 node labels, 36+ relationship tables) |\n| **3b** | `resolve_graph`, `lsp_resolve` | Import/call/impl resolution with confidence scoring + optional LSP deep resolution (rust-analyzer / pyright / typescript-language-server) |\n| **3c** | `cluster_graph`, `get_processes`, `get_impact` | Leiden-class community detection (Louvain + C2 repair) + BFS execution-flow tracing from entry points |\n| **3d** | `search_codebase`, `get_context`, `analyze_codebase`, `detect_changes` | Hybrid BM25 + sparse TF-IDF + RRF search · 360° symbol view · all-in-one analysis · git-diff impact |\n| **4** | `prepare_prd_input` | Bundle verified finding + graph intel → artifact for prd-spec-generator |\n| **6** | `validate_prd_against_graph` | Symbol hallucination · community consistency · process-impact contradiction |\n| **8** | `check_security_gates` | Auth-critical community · unsafe symbol · public-API change · unresolved-import intro · test-coverage gap |\n| **9** | `verify_semantic_diff` | Before/after graph diff with Tarjan SCC cycle detection and regression scoring |\n\n\u003e Stages 5 (PRD generation), 7 (implementation), 10 (benchmark), 11 (deployment), 12 (PR) belong to other systems in the pipeline: [prd-spec-generator](https://github.com/cdeust/prd-spec-generator), the coding agent, CI, and `gh`. This project is the **read-only intelligence** half.\n\n---\n\n## 23 MCP Tools\n\nEvery tool takes structured JSON arguments via the MCP protocol and returns a structured JSON response. No LLM is called from inside any tool — intelligence is the agent's job; the tool's job is safe, fast data movement with invariants.\n\n```\nStage 0:  health_check\nStage 1:  extract_finding · refine_finding\nStage 2:  start_verification · append_clarification · finalize_verification · abort_verification\nStage 3a: index_codebase · query_graph · get_symbol\nStage 3b: resolve_graph · lsp_resolve\nStage 3c: cluster_graph · get_processes · get_impact\nStage 3d: search_codebase · get_context · analyze_codebase · detect_changes\nStage 4:  prepare_prd_input\nStage 6:  validate_prd_against_graph\nStage 8:  check_security_gates\nStage 9:  verify_semantic_diff\n```\n\nEach tool has a JSON Schema enforced at the wire, reason codes on error (no cryptic protocol errors), and a receipt-style response with timing and counts.\n\n---\n\n## Architecture\n\nRust MCP server, hand-rolled stdio JSON-RPC 2.0 (no SDK — we own the wire). Clean Architecture with module boundaries.\n\n```\ntransport (stdio, JSON-RPC framing)\n      ↓\nserver/main.rs  (request dispatch, tool registry)\n      ↓\nhandlers (do_* functions, one per tool)\n      ↓\ncore modules:\n    graph_store        — LadybugDB port (Cypher + UNWIND + prepared statements)\n    parser/{rust,python,typescript,mod}  — tree-sitter AST extractors\n    indexer            — walk + parse + persist pipeline\n    resolver           — cross-file import/call/impl resolution\n    lsp_{client,resolver}  — optional LSP deep resolution\n    clustering         — inline Louvain + C2 repair + process tracing\n    search/{bm25,vector,rrf,mod}  — hybrid search (Tantivy + sparse TF-IDF + RRF)\n    prd_input          — stage 4: bundle for prd-spec-generator\n    prd_validator      — stage 6: validate PRD claims against graph\n    security_gates     — stage 8: auth/unsafe/API/imports/coverage checks\n    semantic_diff      — stage 9: before/after graph regression scoring\n    git_diff           — diff parser + symbol mapping\n```\n\n### Crates\n\nEight crates. Nothing speculative; everything justified.\n\n| Crate | Purpose | License | Why |\n|---|---|---|---|\n| `serde` + `serde_json` | Wire serialization | MIT | JSON-RPC, artifact persistence |\n| `sha2` | Stage-2 transcript digest | MIT | Tamper detection |\n| `lbug` (LadybugDB) | Embedded property graph + Cypher | MIT | Native Cypher, FTS-ready, the Kùzu successor |\n| `tree-sitter` | Incremental parser runtime | MIT | First-class Rust bindings |\n| `tree-sitter-rust` · `-python` · `-typescript` | Language grammars | MIT | Semantic structure without a compiler |\n| `tantivy` | Lucene-grade BM25 | MIT | Real ranked text search, \u003c10ms startup |\n\nDeliberately **not** included: async runtime (we're stdio-blocking), HTTP client, LLM SDK, embedding model runtime (sparse TF-IDF replaces it at zero dep cost).\n\n### Storage\n\nGraphs are per-finding by design (Lamport's isolation invariant): each finding gets its own LadybugDB instance at `\u003coutput_dir\u003e/runs/\u003crun_id\u003e/findings/\u003cfinding_id\u003e/graph/`. Zero-coordination concurrency, trivial cleanup, no cross-finding state leakage. Redundant indexing for shared codebases is acknowledged and mitigated in a later optional cache layer — not shoehorned into the core.\n\n---\n\n## The zetetic standard\n\nInherited from [zetetic-team-subagents](https://github.com/cdeust/zetetic-team-subagents). Not a prompt suggestion — an enforcement rule that holds in code.\n\n| Pillar | Question |\n|---|---|\n| **Logical** | *Is it consistent?* |\n| **Critical** | *Is it true?* |\n| **Rational** | *Is it useful?* |\n| **Essential** | *Is it necessary?* |\n\n**In this codebase it concretely means:**\n\n1. Every algorithm traces to a source. Louvain → *Blondel et al. 2008*. Leiden C2 repair → *Traag et al. 2019*. RRF → *Cormack, Clarke, Büttcher 2009*. SCC → *Tarjan 1972*. BM25 via Tantivy → *Robertson et al. 1994*.\n2. Every named constant has a `// source:` comment. `RRF_K = 60` cites Cormack 2009. `BULK_BATCH_SIZE = 500` cites Kùzu/LadybugDB tuning. `PARSE_TIMEOUT_MICROS = 5_000_000` is justified in the block above it.\n3. No invented numbers. Where a value was chosen by judgment, the comment says so (\"heuristic, not paper-backed\") and cites its operational justification.\n4. Tool responses cite the spec that governs each error reason. `unsafe finding_id (spec §5.1.4, §9.3 Q4): must match [A-Za-z0-9._-]+` — callers see which rule they violated.\n5. When a capability can't be proved at spec time, the tool degrades gracefully and says so in plain language. Example: `lsp_resolve` on a stub binary returns `lsp_probe_failed: found on PATH but didn't respond as an LSP server (stdout closed immediately; likely a stub, proxy, or non-LSP binary)` — not a cryptic protocol error.\n\n---\n\n## Security\n\nFour CRITICAL, four HIGH, three MEDIUM findings were surfaced by a `security-auditor` agent pass and fixed in commit [`512d683`](https://github.com/cdeust/automatised-pipeline/commit/512d683):\n\n- Cypher injection via `insert_edge` → centralized `cypher_str()` escaping (`\\` first, then `'`)\n- Git argument injection → `validate_git_ref` rejects `--`, newlines, NUL; `--` separator before refs\n- Arbitrary binary execution via `lsp_command` → strict allowlist (`rust-analyzer`, `pyright`, `pyright-langserver`, `typescript-language-server`)\n- Symlink traversal → `fs::symlink_metadata` + `MAX_DEPTH`\n- Resource exhaustion → `MAX_FILES=100_000`, `MAX_FILE_BYTES=10 MB`, `MAX_TOTAL_BYTES=2 GB`, `MAX_DEPTH=64`\n- Tree-sitter pathological input → `set_timeout_micros(5_000_000)` + `MAX_PARSE_BYTES=1 MB`\n- `query_graph` read-only → forbidden-keyword whole-word filter (CREATE/DELETE/MERGE/SET/REMOVE/DROP/ALTER/CALL/LOAD)\n- `graph_path` filesystem safety → `validate_graph_path_safe()` before any `remove_dir_all`\n- LSP `rootUri` → RFC 3986 percent-encoding\n- Diff line overflow → `DIFF_LINE_MAX = u64::MAX / 2` guard\n\nEach fix has a test that asserts the exploit is now rejected. Run `cargo test` to see 220 tests pass including the exploit-regression suite.\n\n---\n\n## Scale\n\nVerified by the `dba` agent through compile-and-run probes against lbug 0.15.3:\n\n| Strategy | ms/edge |\n|---|---|\n| Raw string per edge (naive) | 5.36 |\n| Prepared statement, no transaction | 5.48 |\n| `BEGIN TRANSACTION` + prepared + `COMMIT` | 0.70 |\n| **UNWIND + typed `LogicalType::Struct`** | **0.143** |\n\nThe bulk-insert path uses UNWIND with a typed struct schema (the engineer who wrote the first version used `LogicalType::Any` which fails the binder — the typed struct form works). Prepared statements are cached in a `RefCell\u003cHashMap\u003cquery, PreparedStatement\u003e\u003e` on the `GraphStore`. Sparse TF-IDF replaces the dense `N × V × 4B` matrix — **30.5× smaller** on our own codebase (108 KB vs 3.2 MB) and scales linearly with non-zero terms rather than vocab size. Clustering eliminated `probe_node_label_for_process` (per-node Cypher round-trip) in favor of a single in-memory `HashMap\u003cid, label\u003e` population pass.\n\n500-file synthetic Rust fixture indexes in **~38 seconds** end-to-end (parse + resolve + cluster + search index), down from the pre-audit implied \"5 min – 1 hour\" bracket.\n\n---\n\n## Integration with the rest of the stack\n\n```\n                 ┌─────────────────────────────────────────┐\n                 │           Claude Code agent             │\n                 └────────────┬────────────────────────────┘\n                              │ MCP (stdio JSON-RPC)\n                              ↓\n      ┌──────────────────────────────────────────────────┐\n      │             automatised-pipeline                 │  ← this repo\n      │  stage 0 · 1 · 2 · 3a-d · 4 · 6 · 8 · 9          │\n      │  Rust · LadybugDB · tree-sitter · Tantivy        │\n      └──────┬──────────────────┬────────────────────────┘\n             │                  │\n             │                  └────→  stage 5 (PRD gen)\n             │                          [prd-spec-generator]\n             ↓                          TypeScript / Node\n     ┌─────────────────┐                    │\n     │     Cortex      │                    │\n     │  memory engine  │ ←──────────────────┘\n     │  PostgreSQL +   │\n     │    pgvector     │\n     └─────────────────┘\n             ↑\n             │  cross-session memory for findings,\n             │  decisions, lessons learned\n             │\n     ┌─────────────────────────────┐\n     │  zetetic-team-subagents     │\n     │  97 genius + 18 specialists │\n     │  problem-shape routing      │\n     └─────────────────────────────┘\n```\n\n- **Cortex** — every architectural decision made during a pipeline run gets remembered. When the next finding touches a similar area, Cortex surfaces the prior reasoning before you re-derive it.\n- **zetetic-team-subagents** — the genius agents (Shannon, Lamport, Simon, Popper, Feynman, Fermi, dba, architect, security-auditor, engineer) designed this project stage by stage. Every major decision in `stages/*.md` traces to an agent dispatch.\n- **prd-spec-generator** — consumes our `stage-4.prd_input.json` artifact via disk or MCP-to-MCP query of `search_codebase` / `get_context` / `get_impact`. Each in its ideal language: our performance-critical graph work in Rust, their document generation in TypeScript.\n\n---\n\n## Testing\n\n```bash\ncargo test                                          # 220 tests, full suite\ncargo test --release --test scalability_bench       # 500-file synthetic fixture\ncargo test --release --test lbug_bulk_investigation # dba's 9 UNWIND probes\ncargo test --release --test stage3a_integration     # end-to-end per sub-stage\ncargo test --release --test stage9_integration      # before/after diff\ncargo check                                         # zero warnings required\ncargo build --release                               # release binary\n```\n\nEvery stage has an integration test with fixture data. The `lbug_bulk_investigation` test is intentionally preserved — it's the compile-and-run proof that dba's UNWIND pattern works, kept for regression protection and documentation.\n\n---\n\n## Repository layout\n\n```\nautomatised-pipeline/\n├── src/\n│   ├── main.rs                    ← MCP server, 23 tool handlers\n│   ├── tool_schemas.rs            ← JSON Schemas for every tool\n│   ├── lib.rs                     ← re-exports for integration tests\n│   ├── graph_store.rs             ← LadybugDB port (UNWIND + prepared + cached)\n│   ├── parser/\n│   │   ├── mod.rs                 ← language dispatch\n│   │   ├── rust.rs · python.rs · typescript.rs\n│   ├── indexer.rs                 ← walk + parse + persist\n│   ├── resolver.rs                ← cross-file resolution\n│   ├── lsp_client.rs              ← minimal LSP probe + client\n│   ├── lsp_resolver.rs            ← LSP-backed deep resolution\n│   ├── clustering.rs              ← Louvain + C2 repair + BFS process tracing\n│   ├── search/\n│   │   ├── mod.rs                 ← orchestration, get_context, 3-layer qn lookup\n│   │   ├── bm25.rs · vector.rs · rrf.rs\n│   ├── prd_input.rs               ← stage 4\n│   ├── prd_validator.rs           ← stage 6\n│   ├── security_gates.rs          ← stage 8\n│   ├── semantic_diff.rs           ← stage 9\n│   └── git_diff.rs                ← diff parsing + symbol mapping\n├── stages/                        ← locked spec per stage (Shannon, then engineer implements)\n│   ├── stage-1.md · stage-2.md · stage-3.md · stage-3b.md · stage-3c.md\n│   ├── stage-6.md · stage-8.md\n│   ├── stage-1.review.md · stage-3-db-evaluation.md · stage-3-research.md\n│   └── decisions/                 ← Popper / Lamport / Simon verdicts per decision\n├── tests/\n│   ├── stage{3a,3b,3c,3d,4,6,8,9}_integration.rs\n│   ├── multilang_integration.rs\n│   ├── stage3d_hybrid_search.rs\n│   ├── scalability_bench.rs\n│   ├── lbug_bulk_investigation.rs\n│   ├── tfidf_size_report.rs\n│   └── fixtures/multilang/        ← sample.rs · sample.py · sample.ts\n├── .claude/\n│   ├── agents/                    ← 18 specialists + 97 genius agents\n│   ├── skills/ · commands/ · tools/ · hooks/\n│   └── scripts/\n├── .mcp.json\n├── NOTES.md                       ← stages table + growth rule\n├── Cargo.toml\n└── README.md\n```\n\n---\n\n## The zetetic decisions behind the build\n\nEvery major architectural decision was made by a genius agent with a specific problem shape. Stored in `stages/decisions/*.md` and in Cortex.\n\n| Decision | Agent | Verdict |\n|---|---|---|\n| Rust vs C/C++ for the glue layer | **Popper** | Conjecture \"Rust is the right language\" is unfalsified. `lbug` + `tree-sitter` already run native C/C++; Rust is the glue where the borrow checker pays the most. |\n| Graph-per-finding vs graph-per-codebase | **Lamport** | Per-finding. Isolation holds by construction with zero coordination; the redundant-indexing cost is mitigable in an optional cache layer later. |\n| Stage 3a decomposition | **Simon** | Five steps, satisficed against the growth rule; first useful query at step 4. |\n| DB backend choice | **dba** | LadybugDB (`lbug 0.15.3`) — only option simultaneously maintained, native Cypher, embedded, with FTS + vector + algo extensions. |\n| Stage 2 clarification loop shape | **Shannon** | Four-tool state machine with atomic single-file session (no crash window between separate files), unconditional one-round-minimum before finalize. |\n| lbug UNWIND pattern | **dba** | `LogicalType::Struct { fields }` works; `LogicalType::Any` fails the binder — 38× speedup verified by compile-and-run probes. |\n\nAgents are spawned via [zetetic-team-subagents](https://github.com/cdeust/zetetic-team-subagents); each genius is a reasoning pattern (not a persona) with canonical moves and primary-source citations.\n\n---\n\n## Status\n\nPrivate repo by design. Not ready for public release until the full hardening pass is done — security audit fixes are in, correctness fixes are in, scale fixes are in, stages 4/6/8/9 are live, but every capability marked \"live\" above has been verified end-to-end on this machine, not yet in a production context.\n\n**What works today**: indexing Rust / Python / TypeScript codebases end-to-end, resolving cross-file relationships, clustering into communities, tracing processes from entry points, hybrid search, PRD input preparation, PRD claim validation, security gate checking, before/after regression detection.\n\n**What's deferred**:\n- Cross-file indexer batching to unlock the full 38× UNWIND win (currently 1.17× aggregate; per-edge rate is already 0.143 ms)\n- `is_unsafe` extraction in the Rust parser (stage 8 S2 runs in `info`-skip mode pending this)\n- LSP-based deep method resolution on inferred types\n- Multi-repo / workgroup operations (GitNexus `group_*`)\n- Rename / refactor tools (we are read-only by design)\n\n---\n\n## License\n\nMIT — see [LICENSE](LICENSE).\n\n---\n\n\u003cp align=\"center\"\u003e\u003csub\u003eBuilt by \u003ca href=\"https://github.com/cdeust\"\u003ecdeust\u003c/a\u003e. Every stage designed by a genius agent. Every constant sourced.\u003c/sub\u003e\u003c/p\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcdeust%2Fautomatised-pipeline","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fcdeust%2Fautomatised-pipeline","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcdeust%2Fautomatised-pipeline/lists"}