{"id":51792458,"url":"https://github.com/kanutocd/mammoth-search-watch","last_synced_at":"2026-07-20T23:38:14.881Z","repository":{"id":369916342,"uuid":"1291889114","full_name":"kanutocd/mammoth-search-watch","owner":"kanutocd","description":"PostgreSQL-native Search Observability Data Plane built on Mammoth CDC. Implements real-time data sovereignty, RLS multi-tenancy, and layout drift isolation for SERP API payloads with SearchApi as the first supported provider adapter.","archived":false,"fork":false,"pushed_at":"2026-07-07T15:03:15.000Z","size":1136,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-07-20T23:38:10.382Z","etag":null,"topics":["audit-log","cdc","change-data-capture","data-engineering","data-sovereignty","database-replication","llm-ops","multi-tenancy","observability-plane","postgresql","rag-pipeline","realtime","row-level-security","searchapi","serp","serp-api","supabase","wal"],"latest_commit_sha":null,"homepage":"https://kanutocd.github.io/mammoth-search-watch/","language":"Ruby","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/kanutocd.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-07-07T06:48:12.000Z","updated_at":"2026-07-14T03:37:27.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/kanutocd/mammoth-search-watch","commit_stats":null,"previous_names":["kanutocd/mammoth-search-watch"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/kanutocd/mammoth-search-watch","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kanutocd%2Fmammoth-search-watch","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kanutocd%2Fmammoth-search-watch/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kanutocd%2Fmammoth-search-watch/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kanutocd%2Fmammoth-search-watch/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/kanutocd","download_url":"https://codeload.github.com/kanutocd/mammoth-search-watch/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kanutocd%2Fmammoth-search-watch/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35703527,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"ssl_error","status_checked_at":"2026-07-20T02:08:09.736Z","response_time":111,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audit-log","cdc","change-data-capture","data-engineering","data-sovereignty","database-replication","llm-ops","multi-tenancy","observability-plane","postgresql","rag-pipeline","realtime","row-level-security","searchapi","serp","serp-api","supabase","wal"],"created_at":"2026-07-20T23:38:14.256Z","updated_at":"2026-07-20T23:38:14.873Z","avatar_url":"https://github.com/kanutocd.png","language":"Ruby","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003cbr /\u003e\n  \u003cimg src=\"docs/assets/logo/mammoth-horizontal-with-white-background.png\" alt=\"Mammoth PostgreSQL CDC Data Plane\" width=\"750\" style=\"max-width: 100%; height: auto;\"\u003e\n  \u003cbr /\u003e\n  \u003ch1\u003e🐘 Mammoth Search Watch\u003c/h1\u003e\n  \u003cp\u003e\n    \u003cstrong\u003eThe PostgreSQL-Native Search Observability Plane for Enterprise RAG \u0026 LLM Infrastructure\u003c/strong\u003e\n  \u003c/p\u003e\n  \u003cp align=\"center\"\u003e    \n    \u003cimg src=\"https://img.shields.io/badge/postgresql-4169e1\" alt=\"PostgreSQL Core\"\u003e    \n    \u003ca href=\"https://www.searchapi.io\" target=\"_blank\"\u003e \u003cimg src=\"https://img.shields.io/badge/SearchApi-Native%20Support\" alt=\"SearchApi Native Support\"\u003e\u003c/a\u003e\n  \u003c/p\u003e\n\u003c/div\u003e\n\n---\n\n# Mammoth Search Watch\n\n[![Gem Version](https://badge.fury.io/rb/mammoth-search-watch.svg)](https://badge.fury.io/rb/mammoth-search-watch)\n[![CI](https://github.com/kanutocd/mammoth-search-watch/workflows/CI/badge.svg)](https://github.com/kanutocd/mammoth-search-watch/actions)\n[![Ruby Version](https://img.shields.io/badge/ruby-%3E%3D%204.0-ruby.svg)](https://www.ruby-lang.org/en/)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n\n\n## 🌐 Part of the Mammoth Platform\n\n**Mammoth Search Watch** is a concrete production implementation of the open-source **Mammoth Data Plane**. \n\nWhile this data plane remains fully open-source (MIT) for local high-throughput tracking, it natively integrates with the commercial **Mammoth Platform** ecosystem:\n* **Mammoth Control Plane:** Centralized cluster management, global analytics dashboarding, and alerting.\n* **Mammoth Control Agent:** Lightweight telemetry and remote configuration sync.\n* **High-End Extensions:** Advanced predictive drift modeling and automated schema patching.\n\n\n## 🔌 Supported Providers \u0026 Telemetry Adapters\n\n**Mammoth Search Watch** functions as a highly decoupled data plane. Rather than building tightly coupled scrapers, it uses pluggable provider adapters to parse incoming payloads. \n\n* **[SearchApi Platform](https://searchapi.io)** (Production-Ready) — The primary supported data ingestion adapter. Treats incoming live SERP payloads as immutable PostgreSQL facts.\n* **[SearchApi GitHub Org](https://github.com/SearchApi)** — Full compatibility mapping designed to seamlessly pipe search data streams originating from SearchApi's infrastructure tools directly into local CDC pipelines.\n\n\n\n\u003e **PostgreSQL-native Search Observability built on the Mammoth data plane.**\n\n**Observe**. **Persist**. **Prove**.\n\nMammoth Search Watch transforms SERP/search API interactions into durable PostgreSQL facts carried through Mammoth's WAL data plane.\n\nUnlike traditional rank trackers, SEO dashboards, and search monitoring tools, Mammoth Search Watch is designed as a **Search Observability Platform**.\n\nIts philosophy is simple:\n\n\u003e **Observe reality. Persist facts. Deliver changes.**\n\n---\n\n## Why Mammoth Search Watch?\n\nSearch providers tell you **what search results look like now**.\n\nMammoth Search Watch helps you understand:\n\n- **What was searched**\n- **What was returned**\n- **What changed**\n- **When it changed**\n- **and proves it with durable PostgreSQL facts.**\n\n### Designed to enable\n\n- 🔍 Search Observability\n- 📈 SERP Drift Detection\n- 🕒 Historical Search Intelligence\n- 📋 Compliance \u0026 Audit Trails\n- ⚖️ Conflict Resolution\n- 🔄 Replayable Search Events\n- 📊 Search Analytics\n- 🧠 AI-ready Search Data\n- 💰 Search Budget Optimization\n- 🏢 Multi-tenant Operation\n\nThese capabilities are built on a single principle:\n\n\u003e **Observation first. Inference later.**\n\nSearch observations become durable PostgreSQL facts.\n\nEverything else—drift detection, analytics, dashboards, AI, replay, compliance, competitive intelligence, and operational intelligence—is derived from those facts.\n\n---\n\nMammoth Search Watch observes SERP/search API request-response activity, persists normalized observation facts into PostgreSQL, and embeds Mammoth by default to deliver resulting changes through WAL-backed delivery.\n\n**Mammoth Search Watch** is a SERP observation and drift-capture service built on the Mammoth data plane.\n\nIt captures SERP request/response observations, persists them as PostgreSQL facts, and embeds Mammoth so the resulting changes can move through PostgreSQL WAL, replication slots, and reliable downstream delivery.\n\nSearchAPI is the first supported provider adapter.\n\n```text\nTiny observer / adapter\n        ↓\nMammoth Search Watch\n        ↓\nPostgreSQL facts\n        ↓\nPostgreSQL WAL + replication slot\n        ↓\nMammoth\n        ↓\nWebhook / downstream delivery\n```\n\nMammoth Search Watch is intentionally **PostgreSQL-first**, **WAL-centric**, and **operationally boring**. It is not a generic HTTP event bus, not a scraper, and not an SEO dashboard. Sink-only mode is available when you opt out of the embedded Mammoth runtime.\n\n## Status\n\n**Project Status:** Early development.\n\nThe current implementation focuses on establishing the PostgreSQL-first observation pipeline and Mammoth integration.\n\nThe capabilities described in this document represent the long-term vision of Mammoth Search Watch and will be delivered incrementally as the project evolves toward `1.0`.\n\n## Core idea\n\nA SERP API endpoint returns observable search state. Mammoth Search Watch records that state as durable PostgreSQL facts and emits only meaningful changes through Mammoth.\n\n```text\nSERP API endpoint interaction\n        ↓\nrequest observed\nresponse observed\n        ↓\nwatch key + result hash\n        ↓\nPostgreSQL insert\n        ↓\nMammoth delivery\n```\n\n## Who is it for?\n\nMammoth Search Watch is useful for two related audiences:\n\n1. **SERP API company itself** — product telemetry, support evidence, parser regression detection, compliance trails, and drift analytics.\n2. **SERP API company customers** — durable search-change events without rewriting business logic around polling and diffing.\n\n## Architecture boundary\n\nMammoth Search Watch creates PostgreSQL facts. Mammoth operates and delivers them, either embedded by default or as an external service in sink-only mode.\n\n```text\nMammoth Search Watch\n  owns: observation ingestion, normalization, retention, drift fact persistence, embedded Mammoth runtime\n\nMammoth\n  owns: WAL consumption, replication slot handling, delivery, retries, dead letters, health, metrics\n```\n\nThis keeps Mammoth Search Watch true to the Mammoth model:\n\n```text\nPostgreSQL table\n      ↓\nWAL\n      ↓\nreplication slot\n      ↓\nMammoth data plane\n```\n\n## Fragile ingress table\n\nThe first integration boundary is intentionally simple: a fragile, retention-managed PostgreSQL table for observed HTTP activity.\n\nExample shape:\n\n```sql\nCREATE TYPE activity_type AS ENUM ('request', 'response');\n\nCREATE TABLE activities (\n  id BIGSERIAL PRIMARY KEY,\n  tenant_id TEXT NOT NULL,\n  observation_id TEXT NOT NULL,\n  activity_type activity_type NOT NULL,\n  payload JSONB NOT NULL,\n  created_at TIMESTAMPTZ NOT NULL DEFAULT now()\n);\n```\n\nThe `activities` table is an ingress ledger, not canonical long-term history. Rows may be deleted after the configured retention period.\n\nConsumers that want to retain full history may configure a longer retention period or replicate the table elsewhere.\n\n## Multi-tenancy\n\nMammoth Search Watch is designed to be multi-tenant aware.\n\nA PostgreSQL Row-Level Security policy can isolate tenant data:\n\n```sql\nALTER TABLE activities ENABLE ROW LEVEL SECURITY;\n\nCREATE POLICY tenant_isolation_policy\nON activities\nUSING (tenant_id = current_setting('app.current_tenant_id', true))\nWITH CHECK (tenant_id = current_setting('app.current_tenant_id', true));\n```\n\nExample usage:\n\n```sql\nBEGIN;\nSET LOCAL app.current_tenant_id = 'tenant_abc_123';\n\nINSERT INTO activities (\n  tenant_id,\n  observation_id,\n  activity_type,\n  payload\n)\nVALUES (\n  'tenant_abc_123',\n  'obs_123',\n  'request',\n  '{\"engine\":\"google_rank_tracking\",\"q\":\"ruby jobs\"}'::jsonb\n);\n\nCOMMIT;\n```\n\nIf the deployment is not multi-tenant, configure a global tenant id and\nuse it for every insert. The table still stores `tenant_id` on each row;\nthe value just comes from deployment configuration instead of request\ncontext.\n\n## Watch key and result hash\n\nMammoth Search Watch separates request identity from response identity.\n\n```text\nwatch_key / observation_id\n  deterministic identity derived from the normalized request shape\n\nsample_id\n  unique identity for a specific request/response instance\n\nresult_hash\n  deterministic hash of the normalized SERP API result\n```\n\nThe practical deduplication rule is:\n\n```text\nsame watched query + same result hash\n  = no new durable search state\n\nsame watched query + new result hash\n  = new PostgreSQL fact\n  = new WAL event\n  = Mammoth delivery\n```\n\nA derived table may use a uniqueness rule like:\n\n```sql\nUNIQUE (observation_id, result_hash)\n```\n\n## Tiny observer principle\n\nObservers should have a very small footprint.\n\nThey should:\n\n1. copy selected request data,\n2. derive or receive an observation id,\n3. pass the request through,\n4. copy selected response data,\n5. pass the response through unchanged,\n6. enqueue or insert the observation payload.\n\nObservers should not own drift analytics, Mammoth delivery, retention policy, or long-term product behavior.\n\n## Installation\n\nAdd the gem to your application:\n\n```bash\nbundle add mammoth-search-watch\n```\n\nOr install directly:\n\n```bash\ngem install mammoth-search-watch\n```\n\n## Docker\n\nPlanned image:\n\n```text\nghcr.io/kanutocd/mammoth-search-watch:latest\nghcr.io/kanutocd/mammoth-search-watch:v0.1.0\n```\n\nThe default local deployment keeps Mammoth embedded in the Search Watch\nprocess.\n\n```yaml\nservices:\n  postgres:\n    image: postgres:17\n\n  search-watch:\n    image: ghcr.io/kanutocd/mammoth-search-watch:latest\n    environment:\n      DATABASE_URL: postgres://mammoth_search_watch:secret@postgres:5432/search_watch\n      SEARCH_WATCH_CONFIG: /config/search_watch.yml\n      MAMMOTH_CONFIG: /config/mammoth.yml\n    ports:\n      - \"9292:9292\"\n    volumes:\n      - ./config/search_watch.yml:/config/search_watch.yml:ro\n      - ./config/mammoth.yml:/config/mammoth.yml:ro\n      - mammoth_data:/app/.sqlite3\n    depends_on:\n      - postgres\n\nvolumes:\n  mammoth_data:\n  postgres_data:\n```\n\nConcrete deployment manifests are available in [docker-compose.yml](./docker-compose.yml) and the split Kubernetes manifests under `k8s/` folder.\n\n```bash\ndocker compose up -d\nkubectl apply -k k8s\n```\n\n## Kubernetes and Helm\n\nMammoth Search Watch should reuse the Mammoth deployment model where possible.\n\nA dedicated Helm chart should only be introduced when Search Watch needs distinct deployment variants, such as:\n\n- observer ingress service,\n- tenant-specific configuration,\n- retention jobs,\n- RLS/bootstrap migrations,\n- separate service accounts,\n- separate secrets,\n- hosted-vs-customer-pod profiles.\n        \n## Non-goals\n\nMammoth Search Watch is not:\n\n- a browser scraper,\n- a proxy-rotation system,\n- an SEO dashboard,\n- a generic HTTP event bus,\n- a replacement for a SERP API,\n- a replacement for Mammoth.\n\n## Development\n\nAfter checking out the repository:\n\n```bash\nbundle exec bin/setup\nbundle exec rake test\n```\n\nTo bootstrap the PostgreSQL schema:\n\n```bash\nbundle exec mammoth-search-watch bootstrap config/search_watch.yml\n```\n\nTo run retention cleanup:\n\n```bash\nbundle exec mammoth-search-watch retention-cleanup config/search_watch.yml\n```\n\n`mammoth-search-watch start` bootstraps the schema automatically unless\n`lifecycle.bootstrap_on_start` is set to `false`.\n\nIn Docker images, the container entrypoint should call `mammoth-search-watch start`.\n\nFor a long-running periodic cleanup worker, use:\n\n```bash\nbundle exec mammoth-search-watch retention-scheduler config/search_watch.yml\n```\n\nUse the console for local exploration:\n\n```bash\nbundle exec bin/console\n```\n\n## License\n\nThe gem is available as open source under the terms of the [MIT License](https://opensource.org/licenses/MIT).\n\n## Code of Conduct\n\nEveryone interacting with this project is expected to follow the [Code of Conduct](CODE_OF_CONDUCT.md).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkanutocd%2Fmammoth-search-watch","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkanutocd%2Fmammoth-search-watch","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkanutocd%2Fmammoth-search-watch/lists"}