{"id":47015680,"url":"https://github.com/vinaes/md-succ-ai","last_synced_at":"2026-03-11T21:47:50.395Z","repository":{"id":338716644,"uuid":"1157477740","full_name":"vinaes/md-succ-ai","owner":"vinaes","description":"URL to Markdown API — md.succ.ai","archived":false,"fork":false,"pushed_at":"2026-02-25T18:50:57.000Z","size":368,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"master","last_synced_at":"2026-02-25T18:56:38.836Z","etag":null,"topics":["ai-agents","html-to-markdown","markdown","mcp","playwright","rag","readability","url-to-markdown","web-scraping","youtube-transcript"],"latest_commit_sha":null,"homepage":"https://md.succ.ai","language":"JavaScript","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/vinaes.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-02-13T21:36:30.000Z","updated_at":"2026-02-25T18:51:02.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/vinaes/md-succ-ai","commit_stats":null,"previous_names":["vinaes/md-succ-ai"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/vinaes/md-succ-ai","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vinaes%2Fmd-succ-ai","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vinaes%2Fmd-succ-ai/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vinaes%2Fmd-succ-ai/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vinaes%2Fmd-succ-ai/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/vinaes","download_url":"https://codeload.github.com/vinaes/md-succ-ai/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/vinaes%2Fmd-succ-ai/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":30402397,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-03-11T21:02:20.017Z","status":"ssl_error","status_checked_at":"2026-03-11T20:59:32.667Z","response_time":84,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-agents","html-to-markdown","markdown","mcp","playwright","rag","readability","url-to-markdown","web-scraping","youtube-transcript"],"created_at":"2026-03-11T21:47:50.177Z","updated_at":"2026-03-11T21:47:50.388Z","avatar_url":"https://github.com/vinaes.png","language":"JavaScript","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/●%20md.succ.ai-url%20to%20markdown-3fb950?style=for-the-badge\u0026labelColor=0d1117\" alt=\"md.succ.ai\"\u003e\n  \u003cbr/\u003e\u003cbr/\u003e\n  \u003cem\u003eClean Markdown from any URL. Fast, accurate, agent-friendly.\u003c/em\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://md.succ.ai/health\"\u003e\u003cimg src=\"https://img.shields.io/badge/status-live-3fb950?style=flat-square\" alt=\"status\"\u003e\u003c/a\u003e\n  \u003ca href=\"LICENSE\"\u003e\u003cimg src=\"https://img.shields.io/badge/license-FSL--1.1-blue?style=flat-square\" alt=\"license\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://hub.docker.com\"\u003e\u003cimg src=\"https://img.shields.io/badge/docker-node%2022--slim-2496ED?style=flat-square\" alt=\"docker\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://md.succ.ai/docs\"\u003e\u003cimg src=\"https://img.shields.io/badge/docs-OpenAPI-6BA539?style=flat-square\" alt=\"API docs\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"#quick-start\"\u003eQuick Start\u003c/a\u003e •\n  \u003ca href=\"#features\"\u003eFeatures\u003c/a\u003e •\n  \u003ca href=\"#api\"\u003eAPI\u003c/a\u003e •\n  \u003ca href=\"#how-it-works\"\u003eHow It Works\u003c/a\u003e •\n  \u003ca href=\"#self-hosting\"\u003eSelf-Hosting\u003c/a\u003e •\n  \u003ca href=\"#monitoring\"\u003eMonitoring\u003c/a\u003e •\n  \u003ca href=\"#security\"\u003eSecurity\u003c/a\u003e\n\u003c/p\u003e\n\n---\n\n\u003e Convert any webpage, document, feed, or video to clean, readable Markdown. Built for AI agents, MCP tools, and RAG pipelines. Powered by [succ](https://succ.ai).\n\n## Quick Start\n\n```bash\n# Markdown output\ncurl https://md.succ.ai/https://example.com\n\n# JSON output\ncurl -H \"Accept: application/json\" https://md.succ.ai/https://example.com\n\n# Documents (PDF, DOCX, XLSX, CSV)\ncurl https://md.succ.ai/https://example.com/report.pdf\n\n# YouTube transcript\ncurl https://md.succ.ai/https://youtube.com/watch?v=dQw4w9WgXcQ\n\n# RSS/Atom feed\ncurl https://md.succ.ai/https://blog.example.com/feed.xml\n\n# LLM-optimized (30-50% fewer tokens)\ncurl \"https://md.succ.ai/https://example.com?mode=fit\"\n\n# Batch convert\ncurl -X POST https://md.succ.ai/batch \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"urls\": [\"https://example.com\", \"https://httpbin.org/html\"]}'\n```\n\n\u003e **That's it.** No API key, no signup, no SDK. Just prepend `https://md.succ.ai/` to any URL.\n\n## Features\n\n| Feature | Description |\n|---------|-------------|\n| **9-Pass Extraction** | Readability, Defuddle, Article Extractor, CSS selectors, Schema.org, Open Graph, text density, cleaned body — quality-checked at each step |\n| **7 Formats** | HTML, PDF, DOCX, XLSX, CSV, YouTube transcripts, RSS/Atom feeds |\n| **4-Tier Pipeline** | HTTP fetch → headless browser → LLM extraction → BaaS anti-bot bypass |\n| **Batch Conversion** | Convert up to 50 URLs in one request with concurrent processing |\n| **Async + Webhooks** | Submit long conversions and get results via polling or webhook callback |\n| **Structured Extraction** | `/extract` — JSON schema in, structured data out (LLM-powered) |\n| **Quality Scoring** | Each conversion scored 0-1 with A-F grade |\n| **Fit Mode** | LLM-optimized output — pruned boilerplate, 30-50% fewer tokens |\n| **Citation Links** | Numbered references with footer instead of inline links |\n| **Redis Cache** | Two-layer caching (Redis + in-memory fallback), SHA-256 hashed keys |\n| **Rate Limiting** | Per-IP via Redis atomic pipeline, CF-Connecting-IP aware |\n| **Prometheus + Grafana** | 11 custom metrics, pre-provisioned dashboard, auto-scraped |\n| **Structured Logging** | JSON logs via Pino, per-request correlation IDs |\n| **OpenAPI Docs** | Interactive API reference at `/docs` (Scalar UI) |\n\n\u003cdetails\u003e\n\u003csummary\u003eSupported formats\u003c/summary\u003e\n\n| Format | Content-Type | Method |\n|--------|-------------|--------|\n| HTML | `text/html` | 9-pass extraction + Turndown |\n| PDF | `application/pdf` | Text extraction via unpdf |\n| DOCX | `application/vnd...wordprocessingml` | mammoth → HTML → Turndown |\n| XLSX/XLS | `application/vnd...spreadsheetml` | SheetJS → Markdown tables |\n| CSV | `text/csv` | SheetJS → Markdown table |\n| YouTube | `youtube.com`, `youtu.be` | Transcript extraction via innertube API |\n| RSS/Atom | `application/rss+xml`, `application/atom+xml` | Feed parsing with item metadata |\n\nDocuments are also detected by URL extension (`.pdf`, `.docx`, `.xlsx`, `.csv`) when `Content-Type` is `application/octet-stream`.\n\n\u003c/details\u003e\n\n## API\n\n**Base URL:** `https://md.succ.ai`\n**Docs:** [`/docs`](https://md.succ.ai/docs) (interactive Scalar UI) | [`/openapi.json`](https://md.succ.ai/openapi.json) (OpenAPI 3.1 spec)\n\n### Endpoints\n\n| Method | Path | Description |\n|--------|------|-------------|\n| `GET` | `/{url}` | Convert URL to Markdown |\n| `GET` | `/?url={url}` | Same, query param format |\n| `POST` | `/extract` | Structured data extraction via LLM (JSON schema) |\n| `POST` | `/batch` | Batch convert up to 50 URLs |\n| `POST` | `/async` | Async conversion with optional webhook |\n| `GET` | `/job/:id` | Poll async job status |\n| `GET` | `/health` | Health check (includes Redis status) |\n| `GET` | `/docs` | Interactive API reference |\n| `GET` | `/openapi.json` | OpenAPI 3.1 spec |\n\n### Query Parameters\n\n| Parameter | Values | Description |\n|-----------|--------|-------------|\n| `url` | URL | Target URL (alternative to path format) |\n| `links` | `citations` | Convert inline links to numbered references with footer |\n| `mode` | `fit` | Prune boilerplate sections for smaller LLM context |\n| `max_tokens` | number | Truncate output to N tokens (use with `mode=fit`) |\n\n### Response Headers\n\n| Header | Description |\n|--------|-------------|\n| `x-request-id` | Unique request correlation ID |\n| `x-markdown-tokens` | Token count (cl100k_base) |\n| `x-conversion-tier` | `fetch`, `browser`, `baas:scrapfly`, `llm`, `youtube`, `feed`, `document:pdf`, etc. |\n| `x-conversion-time` | Total conversion time in ms |\n| `x-extraction-method` | Extraction pass used (`readability`, `defuddle`, `browser-raw`, etc.) |\n| `x-quality-score` | Quality score 0-1 |\n| `x-quality-grade` | Quality grade A-F |\n| `x-readability` | `true` if Readability extracted clean content |\n| `x-cache` | `hit` or `miss` (Redis-backed) |\n| `x-ratelimit-limit` | Max requests per window |\n| `x-ratelimit-remaining` | Requests remaining in current window |\n| `x-ratelimit-reset` | Window reset timestamp (Unix seconds) |\n\n### Rate Limits\n\n| Endpoint | Limit |\n|----------|-------|\n| `GET /*` | 60 req/min per IP |\n| `POST /extract` | 10 req/min per IP |\n| `POST /batch` | 5 req/min per IP |\n| `POST /async` | 10 req/min per IP |\n\n\u003cdetails\u003e\n\u003csummary\u003eJSON response format\u003c/summary\u003e\n\n```json\n{\n  \"title\": \"Example Domain\",\n  \"url\": \"https://example.com\",\n  \"content\": \"# Example Domain\\n\\nThis domain is for use in...\",\n  \"fit_markdown\": \"# Example Domain\\n\\nThis domain is...\",\n  \"fit_tokens\": 20,\n  \"excerpt\": \"This domain is for use in documentation examples...\",\n  \"tokens\": 33,\n  \"tier\": \"fetch\",\n  \"readability\": true,\n  \"method\": \"readability\",\n  \"quality\": { \"score\": 0.85, \"grade\": \"A\" },\n  \"time_ms\": 245\n}\n```\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eBatch conversion\u003c/summary\u003e\n\n```bash\ncurl -X POST https://md.succ.ai/batch \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"urls\": [\n      \"https://example.com\",\n      \"https://httpbin.org/html\",\n      \"https://github.com\"\n    ],\n    \"options\": {\n      \"mode\": \"fit\",\n      \"links\": \"citations\"\n    }\n  }'\n```\n\nReturns an array of results. Up to 50 URLs, processed with 10-way concurrency. Per-URL 60s timeout.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eAsync conversion with webhook\u003c/summary\u003e\n\n```bash\n# Submit async job\ncurl -X POST https://md.succ.ai/async \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"url\": \"https://example.com\",\n    \"callback_url\": \"https://your-server.com/webhook\"\n  }'\n# → {\"job_id\": \"abc12345\", \"status\": \"processing\", \"poll_url\": \"/job/abc12345\"}\n\n# Poll for result\ncurl https://md.succ.ai/job/abc12345\n```\n\nWebhook delivers JSON `POST` to `callback_url` on completion/failure. HTTPS required, 3 retries with exponential backoff. Private/internal addresses blocked (SSRF-safe).\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eStructured data extraction\u003c/summary\u003e\n\n```bash\ncurl -X POST https://md.succ.ai/extract \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"url\": \"https://github.com/trending\",\n    \"schema\": {\n      \"type\": \"object\",\n      \"properties\": {\n        \"repositories\": {\n          \"type\": \"array\",\n          \"items\": {\n            \"type\": \"object\",\n            \"properties\": {\n              \"name\": { \"type\": \"string\" },\n              \"author\": { \"type\": \"string\" },\n              \"description\": { \"type\": \"string\" },\n              \"stars_today\": { \"type\": \"number\" }\n            }\n          }\n        }\n      }\n    }\n  }'\n```\n\nReturns structured JSON matching the provided schema, extracted by LLM. Automatically retries with headless browser for SPA/JS-heavy sites when initial extraction returns empty data.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eMore examples\u003c/summary\u003e\n\n```bash\n# Citation-style links (numbered references)\ncurl \"https://md.succ.ai/?url=https://en.wikipedia.org/wiki/Markdown\u0026links=citations\"\n\n# LLM-optimized output (pruned boilerplate)\ncurl \"https://md.succ.ai/?url=https://htmx.org/docs/\u0026mode=fit\"\n\n# Token limit\ncurl \"https://md.succ.ai/?url=https://example.com\u0026mode=fit\u0026max_tokens=4000\"\n\n# RSS feed as markdown\ncurl https://md.succ.ai/https://hnrss.org/frontpage\n```\n\n\u003c/details\u003e\n\n## How It Works\n\n4-tier conversion pipeline — each tier only activates if the previous one produced insufficient quality:\n\n```\nURL ──→ Cache hit? ──→ Return cached result (Redis, dynamic TTL)\n         │\n         ├─ YouTube? ──→ Transcript extraction (innertube API)\n         │\n         ├─ RSS/Atom feed? ──→ Feed parsing with item metadata\n         │\n         ├─ Document? (PDF, DOCX, XLSX, CSV)\n         │   └─→ Document converter → Markdown\n         │\n         ├─ Tier 1: HTTP fetch + 9-pass extraction\n         │   └─→ Readability → Defuddle → Article Extractor → CSS selectors\n         │       → Schema.org → Open Graph → Text density → Body fallback\n         │\n         ├─ Tier 2: Camoufox headless browser (SPA/JS-heavy)\n         │   └─→ Same 9-pass pipeline on rendered DOM\n         │\n         ├─ Tier 2.5: LLM extraction (quality \u003c B)\n         │   └─→ nano-gpt API → content extraction\n         │\n         └─ Tier 3: BaaS anti-bot bypass (CF Turnstile / quality \u003c D)\n             └─→ ScrapFly → ZenRows → ScrapingBee (rotation)\n             └─→ Same 9-pass pipeline on returned HTML\n```\n\nCloudflare challenge pages are detected automatically. When fetch gets a CF challenge, browser is skipped (saves IP), and BaaS providers handle the bypass.\n\nWhen both LLM and BaaS are needed, they race in parallel — saves 30-45s vs sequential.\n\n\u003cdetails\u003e\n\u003csummary\u003eCaching\u003c/summary\u003e\n\nTwo-layer cache system backed by Redis 7:\n\n| Content | TTL | Key |\n|---------|-----|-----|\n| HTML pages | 5 min | `cache:{sha256(url+options)}` |\n| Browser renders | 10 min | Same |\n| YouTube transcripts | 1 hr | Same |\n| Documents | 2 hr | Same |\n| /extract results | 1 hr | `extract:{sha256(url)}:{sha256(schema)}` |\n\nCache keys use SHA-256 hashes to prevent poisoning via long/malicious URLs. Tracking parameters (UTM, fbclid, gclid, etc.) are stripped before hashing. Falls back to in-memory Map when Redis is unavailable.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eStack\u003c/summary\u003e\n\n| Component | Role |\n|-----------|------|\n| [Hono](https://hono.dev) | HTTP framework |\n| [Pino](https://getpino.io) | Structured JSON logging |\n| [Mozilla Readability](https://github.com/mozilla/readability) | Primary content extraction |\n| [Defuddle](https://github.com/nicedoc/defuddle) | Obsidian team's content extraction |\n| [@extractus/article-extractor](https://github.com/nicedoc/extractus) | Alternative extraction heuristics |\n| [Turndown](https://github.com/mixmark-io/turndown) | HTML → Markdown conversion |\n| [linkedom](https://github.com/WebReflection/linkedom) | Lightweight DOM parser |\n| [Camoufox](https://github.com/daijro/camoufox) | Firefox fork with C++ anti-detection |\n| [Redis](https://redis.io) + [ioredis](https://github.com/redis/ioredis) | Cache, rate limiting, job storage |\n| [prom-client](https://github.com/siimon/prom-client) | Prometheus metrics |\n| [unpdf](https://github.com/unjs/unpdf) | PDF text extraction |\n| [mammoth](https://github.com/mwilliamson/mammoth.js) | DOCX → HTML conversion |\n| [SheetJS](https://sheetjs.com) | XLSX/XLS/CSV parsing |\n| [NanoGPT](https://nano-gpt.com) | LLM API for Tier 2.5 and /extract |\n| [Ajv](https://ajv.js.org) | JSON Schema validation for /extract |\n| [gpt-tokenizer](https://github.com/niieani/gpt-tokenizer) | cl100k_base token counting |\n| [nanoid](https://github.com/ai/nanoid) | Request/job IDs |\n\n\u003c/details\u003e\n\n## Self-Hosting\n\n### Docker (recommended)\n\n```bash\ngit clone https://github.com/vinaes/md-succ-ai.git\ncd md-succ-ai\ncp .env.example .env  # edit with your API keys and passwords\ndocker compose up -d\n```\n\nThis starts four containers:\n\n| Container | Purpose | Port |\n|-----------|---------|------|\n| **md-succ-ai** | API server with Camoufox browser fallback | 127.0.0.1:3100 |\n| **md-succ-redis** | Redis 7 (cache, rate limiting, jobs) | internal |\n| **md-succ-prometheus** | Prometheus metrics collector | internal |\n| **md-succ-grafana** | Grafana dashboards | 127.0.0.1:3200 |\n\nThe API is available at `http://localhost:3100`.\n\n### Local (without Docker)\n\n```bash\nnpm install\nnpx camoufox-js fetch\nnpm start\n```\n\n\u003e Redis is optional for local development. Without Redis, caching and rate limiting fall back to in-memory Map, and async jobs are unavailable.\n\n\u003cdetails\u003e\n\u003csummary\u003eEnvironment variables\u003c/summary\u003e\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `PORT` | `3000` | Server port |\n| `ENABLE_BROWSER` | `true` | Enable Camoufox browser fallback |\n| `NODE_ENV` | `production` | Node environment |\n| `REDIS_URL` | `redis://redis:6379` | Redis connection URL (with password in Docker) |\n| `REDIS_PASSWORD` | — | Redis authentication password (required in Docker) |\n| `GRAFANA_PASSWORD` | — | Grafana admin password (required in Docker) |\n| `NANOGPT_API_KEY` | — | nano-gpt API key for LLM tier and /extract |\n| `NANOGPT_MODEL` | `meta-llama/llama-3.3-70b-instruct` | LLM model for content extraction (Tier 2.5) |\n| `NANOGPT_EXTRACT_MODEL` | same as `NANOGPT_MODEL` | LLM model for `/extract` endpoint |\n| `SCRAPFLY_API_KEY` | — | [ScrapFly](https://scrapfly.io) anti-bot bypass (1000 credits/mo free) |\n| `ZENROWS_API_KEY` | — | [ZenRows](https://zenrows.com) anti-bot bypass (1000 credits trial) |\n| `SCRAPINGBEE_API_KEY` | — | [ScrapingBee](https://scrapingbee.com) anti-bot bypass (1000 credits one-time) |\n\nBaaS providers are optional. When configured, they activate as Tier 3 for Cloudflare-protected sites. Providers are tried in order; if one hits rate limits, the next is used automatically.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003eNginx reverse proxy\u003c/summary\u003e\n\nAn example nginx config is in `nginx/md.succ.ai.conf`:\n\n- Rate limiting: 10 req/s per IP, burst 20\n- Connection limit: 10 concurrent per IP\n- Proxy timeouts: 60s read (for browser renders)\n- POST endpoints with appropriate body limits\n- HSTS, security headers (nosniff, X-Frame-Options, Referrer-Policy)\n- `/metrics` blocked (403)\n- `/grafana/` proxied to Grafana container with WebSocket support\n\n\u003c/details\u003e\n\n## Monitoring\n\nThe project ships with a full Prometheus + Grafana stack:\n\n**Prometheus** scrapes the `/metrics` endpoint every 10s (internal Docker network only).\n\n**Grafana** is pre-provisioned with a 15-panel dashboard:\n\n- Request rate, response time percentiles (p50/p95/p99)\n- Conversion tier distribution, cache hit rate\n- Quality score distribution, tokens per conversion\n- Rate limit rejections, async job status\n- Browser pool utilization, webhook deliveries\n- Node.js process metrics (CPU, memory, event loop lag)\n\nAccess Grafana at `https://your-domain/grafana/` (proxied via nginx).\n\n### Custom Metrics\n\n| Metric | Type | Labels |\n|--------|------|--------|\n| `http_requests_total` | Counter | method, route, status |\n| `http_request_duration_seconds` | Histogram | method, route, status |\n| `conversion_tier_total` | Counter | tier |\n| `conversion_tokens` | Histogram | tier |\n| `conversion_quality` | Histogram | tier |\n| `cache_hits_total` | Counter | source |\n| `cache_misses_total` | Counter | — |\n| `rate_limit_rejections_total` | Counter | route |\n| `browser_pool_active` | Gauge | — |\n| `async_jobs_total` | Counter | status |\n| `webhook_deliveries_total` | Counter | status |\n\nPlus Node.js default metrics (CPU, memory, event loop, GC) via `prom-client`.\n\n## Security\n\n- **SSRF protection** — URL validation, DNS resolution checks (IPv4 + IPv6), redirect validation per hop, Camoufox route blocking, webhook callback DNS validation\n- **Private IP blocking** — 127/8, 10/8, 172.16/12, 192.168/16, 169.254/16, CGNAT, cloud metadata hostnames, hex/octal IP formats, IPv6 mapped addresses\n- **Input limits** — 5MB response size, 5 max redirects, content-type validation, body size limits per endpoint\n- **Output sanitization** — Error messages stripped of internal paths/stack traces, URLs sanitized in responses\n- **Cache security** — SHA-256 hashed keys (no URL poisoning), tracking params stripped, Redis LRU eviction (128MB cap)\n- **Redis authentication** — `--requirepass` with password from .env, authenticated connection URL\n- **API key safety** — BaaS API keys only used in outbound requests, never logged or exposed in responses\n- **LLM hardening** — Prompt injection protection (HTML sanitization, document delimiters, output validation), schema field whitelist, blocked schema keywords ($ref, $defs, etc.)\n- **Rate limiting** — Per-IP via Redis INCR+EXPIRE (atomic pipeline), CF-Connecting-IP support, in-memory fallback\n- **Security headers** — HSTS, X-Content-Type-Options, X-Frame-Options, Referrer-Policy, Permissions-Policy\n- **CDN integrity** — Subresource Integrity (SRI) on third-party scripts\n- **Container security** — Non-root user (`mduser`), `no-new-privileges`, pinned image versions\n- **CF challenge detection** — Cloudflare challenge pages detected and handled without wasting browser/BaaS credits\n\n## Architecture\n\n```\n                    ┌──────────────────┐\n                    │   Cloudflare     │\n                    │   (TLS + CDN)    │\n                    └────────┬─────────┘\n                             │\n                    ┌────────▼─────────┐\n                    │   nginx          │\n                    │   (rate limit,   │\n                    │    HSTS, proxy)  │\n                    └────────┬─────────┘\n                             │\n         ┌───────────────────┼───────────────────┐\n         │                   │                   │\n┌────────▼─────────┐ ┌──────▼───────┐ ┌─────────▼────────┐\n│  md-succ-ai      │ │  Prometheus  │ │  Grafana         │\n│  (Node 22, Hono) │ │  (scrape     │ │  (dashboards,    │\n│  Camoufox       │ │   /metrics)  │ │   alerting)      │\n│  BaaS clients    │ └──────────────┘ └──────────────────┘\n│  Pino logging    │\n└────────┬─────────┘\n         │\n┌────────▼─────────┐\n│  Redis 7         │\n│  (cache, rate    │\n│   limit, jobs)   │\n└──────────────────┘\n```\n\n## License\n\n[FSL-1.1-Apache-2.0](LICENSE) — Free for non-competitive use. Apache 2.0 after 2 years.\n\n\u003e **Disclaimer:** Not affiliated with [NanoGPT](https://nano-gpt.com). LLM features use the NanoGPT API for pay-per-prompt model access.\n\n---\n\nPart of the [succ](https://succ.ai) ecosystem.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fvinaes%2Fmd-succ-ai","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fvinaes%2Fmd-succ-ai","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fvinaes%2Fmd-succ-ai/lists"}