{"id":51335439,"url":"https://github.com/kinorai/omnifeed","last_synced_at":"2026-07-02T02:07:47.984Z","repository":{"id":361624873,"uuid":"1255190289","full_name":"kinorai/omnifeed","owner":"kinorai","description":"LLM-friendly web crawler \u0026 scraper with a dedicated Reddit engine, built on Crawl4AI — Open WebUI compatible","archived":false,"fork":false,"pushed_at":"2026-06-29T16:58:02.000Z","size":468,"stargazers_count":4,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-06-29T18:25:55.002Z","etag":null,"topics":["ai","ai-agents","claude","codex","crawl4ai","firecrawl-alternative","html-to-markdown","llm","mcp","mcp-server","model-context-protocol","open-webui","opencode","rag","reddit","searxng","self-hosted","web-crawler","web-scraping","web-search"],"latest_commit_sha":null,"homepage":"https://kinorai.github.io/omnifeed/","language":"Go","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/kinorai.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","dco":null,"cla":null}},"created_at":"2026-05-31T14:19:34.000Z","updated_at":"2026-06-29T16:58:06.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/kinorai/omnifeed","commit_stats":null,"previous_names":["kinorai/crawl4ai-reddit-proxy","kinorai/omnifeed"],"tags_count":21,"template":false,"template_full_name":null,"purl":"pkg:github/kinorai/omnifeed","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kinorai%2Fomnifeed","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kinorai%2Fomnifeed/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kinorai%2Fomnifeed/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kinorai%2Fomnifeed/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/kinorai","download_url":"https://codeload.github.com/kinorai/omnifeed/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kinorai%2Fomnifeed/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35029805,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-02T02:00:06.368Z","response_time":173,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai","ai-agents","claude","codex","crawl4ai","firecrawl-alternative","html-to-markdown","llm","mcp","mcp-server","model-context-protocol","open-webui","opencode","rag","reddit","searxng","self-hosted","web-crawler","web-scraping","web-search"],"created_at":"2026-07-02T02:07:47.263Z","updated_at":"2026-07-02T02:07:47.972Z","avatar_url":"https://github.com/kinorai.png","language":"Go","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003c!-- markdownlint-disable MD033 MD041 --\u003e\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://capsule-render.vercel.app/api?type=waving\u0026color=0:FF4500,100:7C3AED\u0026height=220\u0026section=header\u0026text=omnifeed\u0026fontSize=82\u0026fontColor=ffffff\u0026animation=fadeIn\u0026fontAlignY=36\" alt=\"omnifeed\" width=\"100%\"/\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://readme-typing-svg.demolab.com?font=Fira+Code\u0026weight=700\u0026size=28\u0026color=FF4500\u0026center=true\u0026vCenter=true\u0026multiline=true\u0026repeat=false\u0026duration=1500\u0026pause=500\u0026width=860\u0026height=110\u0026lines=Self-hosted+web+search+%2B+LLM-friendly+crawling;with+dedicated+Reddit+and+Hacker+News+engines\" alt=\"Self-hosted web search + LLM-friendly crawling, with dedicated Reddit and Hacker News engines\"/\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://github.com/kinorai/omnifeed/actions/workflows/ci.yml\"\u003e\u003cimg src=\"https://img.shields.io/github/actions/workflow/status/kinorai/omnifeed/ci.yml?branch=main\u0026label=CI\u0026style=flat-square\" alt=\"CI\"/\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/kinorai/omnifeed/releases\"\u003e\u003cimg src=\"https://img.shields.io/github/v/release/kinorai/omnifeed?style=flat-square\u0026color=FF4500\" alt=\"Release\"/\u003e\u003c/a\u003e\n  \u003ca href=\"https://hub.docker.com/r/kinorai/omnifeed\"\u003e\u003cimg src=\"https://img.shields.io/docker/pulls/kinorai/omnifeed?style=flat-square\u0026logo=docker\u0026logoColor=white\u0026color=2496ED\" alt=\"Docker pulls\"/\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\nomnifeed gives an AI agent the full research loop — \u003cb\u003esearch → URLs → content\u003c/b\u003e — against self-hosted\n\u003ca href=\"https://github.com/searxng/searxng\"\u003eSearXNG\u003c/a\u003e + \u003ca href=\"https://github.com/unclecode/crawl4ai\"\u003ecrawl4ai\u003c/a\u003e,\nwith a dedicated Reddit engine that returns \u003cb\u003efull comment trees as \u003ca href=\"https://github.com/toon-format/toon\"\u003eTOON\u003c/a\u003e\u003c/b\u003e\n(~40% fewer tokens than JSON, lossless) and \u003cb\u003eno Reddit API key\u003c/b\u003e.\n\u003c/p\u003e\n\n- **`web_search`** queries a SearXNG instance (Google/Bing/DDG, Reddit included) and returns ranked URLs with titles and snippets.\n- **`fetch_url`** renders any URL through crawl4ai as clean markdown — and Reddit URLs (threads *and* `/r/{sub}` listings) come back as TOON, as do Hacker News threads and front-page / Ask / Show feeds (read directly from the Algolia HN API).\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"100%\"\u003e\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Activities/Sparkles.png\" width=\"26\" height=\"26\" /\u003e Why omnifeed\n\n| | omnifeed | Other Reddit MCPs |\n|---|---|---|\n| Web search → crawl in one self-hosted service | ✅ SearXNG + crawl4ai | ❌ search-only or crawl-only |\n| Full comment tree (`/api/morechildren` expansion) | ✅ up to 40 rounds (~4k comments) | ❌ |\n| Token-efficient output | ✅ TOON, ~40% smaller than JSON | ❌ verbose JSON or truncated bodies |\n| Generic crawl fallback for non-Reddit URLs | ✅ via crawl4ai | ❌ |\n| Front-ends | MCP **+** Open WebUI **+** REST | MCP only (most) |\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"100%\"\u003e\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Clapper%20Board.png\" width=\"26\" height=\"26\" /\u003e Demo\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"assets/demo.gif\" alt=\"omnifeed demo — web_search then fetch_url returning a Reddit thread as TOON\" width=\"100%\"/\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\u003csub\u003e\u003ccode\u003eweb_search\u003c/code\u003e → pick a URL → \u003ccode\u003efetch_url\u003c/code\u003e → full Reddit comment tree as TOON.\u003c/sub\u003e\u003c/p\u003e\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"100%\"\u003e\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Travel%20and%20places/Rocket.png\" width=\"26\" height=\"26\" /\u003e Quick start\n\n```bash\n# Fetch the compose file + SearXNG settings, then start:\ncurl -fsSL https://raw.githubusercontent.com/kinorai/omnifeed/main/docker-compose.yml -o docker-compose.yml\ncurl -fsSL --create-dirs https://raw.githubusercontent.com/kinorai/omnifeed/main/searxng/settings.yml -o searxng/settings.yml\ndocker compose up\n```\n\nStarts omnifeed + SearXNG + crawl4ai — **tokenless out of the box** (the compose file sets `OMNIFEED_DEV_NO_AUTH=true`), so `docker compose up` just works. Point Open WebUI at `http://localhost:8080` with `WEB_LOADER_ENGINE=external`. (SearXNG is mounted with `searxng/settings.yml`, which enables the `json` format `web_search` needs.) See **Authentication** below to require a bearer token.\n\n### \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Electric%20Plug.png\" width=\"22\" height=\"22\" /\u003e As an MCP server\n\nWorks with any MCP client — **Claude Code, Cursor, Codex, Gemini CLI, OpenCode, Windsurf, Pi**, and more.\n\n**HTTP — recommended.** `docker compose up` already runs the MCP server on `:8081` (tokenless in dev mode), so the simplest setup is no extra container at all — point your client at the URL:\n\n```jsonc\n{ \"mcpServers\": { \"omnifeed\": { \"url\": \"http://localhost:8081/mcp\" } } }\n```\n\n**Stdio — for clients that only speak stdio.** A stdio server is spawned and owned by your client (it pipes JSON-RPC over the process's stdin/stdout), so it can't be a long-running compose service — but you can launch the `mcp` profile from this compose file, which keeps every setting (upstreams, network, image) in one place:\n\n```jsonc\n{\n  \"mcpServers\": {\n    \"omnifeed\": {\n      \"command\": \"docker\",\n      \"args\": [\"compose\", \"-f\", \"/abs/path/to/docker-compose.yml\", \"run\", \"-T\", \"--rm\", \"mcp\"]\n    }\n  }\n}\n```\n\n`run -T` disables the TTY so JSON-RPC pipes cleanly; the container joins the stack's network and reuses crawl4ai/SearXNG. Bring the stack up first (`docker compose up -d`) so the upstreams are healthy.\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eStandalone stdio — without the compose stack\u003c/b\u003e\u003c/summary\u003e\n\nSpawn the container directly and tell it where crawl4ai/SearXNG are reachable (omnifeed exits at startup without `OMNIFEED_CRAWL4AI_URL`):\n\n```jsonc\n{\n  \"mcpServers\": {\n    \"omnifeed\": {\n      \"command\": \"docker\",\n      \"args\": [\n        \"run\", \"--rm\", \"-i\",\n        \"-e\", \"OMNIFEED_CRAWL4AI_URL=http://host.docker.internal:11235/crawl\",\n        \"-e\", \"OMNIFEED_SEARXNG_URL=http://host.docker.internal:8080\",\n        \"kinorai/omnifeed:latest\", \"--mcp-stdio\"\n      ]\n    }\n  }\n}\n```\n\nOn Linux, add `\"--add-host=host.docker.internal:host-gateway\"` to the args so `host.docker.internal` resolves.\n\u003c/details\u003e\n\nTools: **`fetch_url`** (always) and **`web_search`** (only when `OMNIFEED_SEARXNG_URL` is set). The intended loop is `web_search` → pick URLs → `fetch_url`.\n\n`/crawl` returns `[{\"page_content\": \"...\", \"metadata\": {...}}]` — already the shape of a LangChain / LlamaIndex `Document`, so wrapping it as a custom document loader takes only a few lines.\n\n### \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Locked%20with%20Key.png\" width=\"22\" height=\"22\" /\u003e Authentication\n\nThe compose stack runs **tokenless** for local use (`OMNIFEED_DEV_NO_AUTH=true`). To require a bearer token instead, generate one — this is the value clients send as `Authorization: Bearer \u003ctoken\u003e`, so copy it:\n\n```bash\nopenssl rand -hex 32        # ← your token; copy this\n```\n\nThen in `docker-compose.yml` set `OMNIFEED_API_KEY` to that value and remove `OMNIFEED_DEV_NO_AUTH`. Without a key (and without `OMNIFEED_DEV_NO_AUTH=true`) the proxy **refuses to start**, so it can't be left open by accident. Stdio MCP needs no token — it inherits the trust of the process that spawned it.\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"100%\"\u003e\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Gear.png\" width=\"26\" height=\"26\" /\u003e Configuration\n\nEverything is configured with `OMNIFEED_`-prefixed environment variables. In practice you only ever **set three** — `OMNIFEED_API_KEY`, `OMNIFEED_CRAWL4AI_URL`, and (optionally) `OMNIFEED_SEARXNG_URL`. The rest have sane defaults.\n\n| Variable | Default | Purpose |\n|---|---|---|\n| `OMNIFEED_API_KEY` | _(unset)_ | Bearer token for `/crawl`, `/search`, `/mcp`. If unset, the proxy refuses to start unless `OMNIFEED_DEV_NO_AUTH=true`. Stdio MCP is unaffected. |\n| `OMNIFEED_CRAWL4AI_URL` | _(required)_ | Upstream crawl4ai endpoint. Reddit + the generic fallback fetch through it (the Hacker News engine reads `hn.algolia.com` directly); if empty, the proxy exits at startup. |\n| `OMNIFEED_SEARXNG_URL` | _(unset)_ | Upstream SearXNG base URL (e.g. `http://searxng:8080`). When unset, `web_search` / `/search` are not exposed. The instance must enable the `json` format. |\n| `OMNIFEED_DEV_NO_AUTH` | `false` | Run the HTTP transports with **no** auth when no key is set (local/dev only). Ignored if a key is set. |\n| `OMNIFEED_LISTEN_ADDR` | `:8080` | HTTP listen address (`/crawl`, `/search`) |\n| `OMNIFEED_MCP_LISTEN_ADDR` | `:8081` | MCP HTTP/SSE listen address |\n| `OMNIFEED_MCP_STDIO` | `false` | Run MCP over stdio (also via `--mcp-stdio`) |\n| `OMNIFEED_METRICS_ADDR` | `:9090` | Prometheus + health listen address |\n| `OMNIFEED_CRAWL4AI_TIMEOUT` | `90s` | Per-call timeout to crawl4ai |\n| `OMNIFEED_CRAWL4AI_KEEP_LINKS` | `true` | Keep hyperlink anchor text + external links in fetched markdown. Set `false` for leaner, link-stripped output (loses link-dense content like HN titles). |\n| `OMNIFEED_CRAWL4AI_PRUNE_THRESHOLD` | `0.48` | PruningContentFilter cutoff (0–1) for the generic engine. Raise it to strip more boilerplate/duplicated chrome from noisy pages; lower it to keep more. |\n| `OMNIFEED_CRAWL4AI_WAIT_UNTIL` | `domcontentloaded` | crawl4ai page-ready signal (`domcontentloaded` \\| `load` \\| `networkidle` \\| `commit`). `domcontentloaded` fires before client-side frameworks hydrate, so JS-only SPAs can render empty; set `networkidle` to wait for them (slower on every page). |\n| `OMNIFEED_SEARXNG_TIMEOUT` | `15s` | Per-query timeout to SearXNG |\n| `OMNIFEED_SEARCH_MAX_RESULTS` | `25` | Hard cap on the search `limit` argument (1–100) |\n| `OMNIFEED_REDDIT_TIMEOUT` | `4m` | Wall-clock cap for a Reddit thread expansion |\n| `OMNIFEED_REDDIT_MAX_ROUNDS` | `3` | Default `/api/morechildren` rounds (max 40 via `?expand=full`) |\n| `OMNIFEED_REDDIT_FORMAT` | `toon` | Default Reddit output: `toon` or `json` |\n| `OMNIFEED_REDDIT_FETCH_LIMIT` | `500` | Reddit `limit`: max comments fetched in the initial tree |\n| `OMNIFEED_REDDIT_DEPTH` | `20` | Reddit `depth`: max nesting depth of the initial tree |\n| `OMNIFEED_REDDIT_SORT` | `top` | Reddit `sort`: one of `confidence` (=best), `top`, `new`, `controversial`, `old`, `random`, `qa`, `live` |\n| `OMNIFEED_REDDIT_MAX_COMMENTS` | `0` | Hard cap on total comments emitted after expansion (0 = unlimited) |\n| `OMNIFEED_REDDIT_MAX_TOP_LEVEL` | `0` | Hard cap on top-level comment threads, replies included (0 = unlimited) |\n| `OMNIFEED_REDDIT_KEEP_CREATED` | `true` | Include each comment's `created` timestamp |\n| `OMNIFEED_REDDIT_KEEP_DEPTH` | `false` | Include each comment's `depth` field |\n| `OMNIFEED_MAX_URLS_PER_REQUEST` | `30` | Cap on `urls[]` length |\n| `OMNIFEED_PER_DOMAIN_CONCURRENCY` | `2` | Max concurrent requests to one domain |\n| `OMNIFEED_PER_DOMAIN_DELAY` | `1500ms` | Minimum delay between same-domain requests |\n| `OMNIFEED_BLOCK_PRIVATE_IPS` | `true` | SSRF protection (keep on in production) |\n| `OMNIFEED_LOG_LEVEL` | `info` | `debug`/`info`/`warn`/`error` |\n| `OMNIFEED_LOG_FORMAT` | `json` | `json` or `text` |\n| `OMNIFEED_ENABLE_PPROF` | `false` | Expose `/debug/pprof/*` (opt-in) |\n\n### Controlling Reddit response size\n\nA Reddit thread's comment tree can be huge. The size knobs come in two kinds — it matters which is which:\n\n- **Upstream Reddit params** — forwarded verbatim to Reddit's API, so Reddit owns their behavior: `OMNIFEED_REDDIT_FETCH_LIMIT` → `limit`, `OMNIFEED_REDDIT_DEPTH` → `depth`, `OMNIFEED_REDDIT_SORT` → `sort`. They shape *what Reddit sends back* (less latency, fewer tokens) but are **approximate**, and `limit`/`depth` bound only the **initial** fetch. Semantics are Reddit's, not ours — see \u003chttps://www.reddit.com/dev/api/\u003e → `GET [/r/subreddit]/comments/article` (`limit` = \"maximum number of comments to return\", `depth` = \"maximum depth of subtrees\").\n- **omnifeed engine caps** — our own, applied *after* fetch + expansion, so they're **exact and independent of Reddit**: `OMNIFEED_REDDIT_MAX_COMMENTS` (truncate the flat comment list) and `OMNIFEED_REDDIT_MAX_TOP_LEVEL` (keep the first N top-level threads, in `sort` order, with their replies).\n\nRule of thumb: reach for the **upstream params** to fetch less from Reddit; reach for the **engine caps** when you need a guaranteed ceiling — `OMNIFEED_REDDIT_MAX_ROUNDS` expansion adds comments on top of `limit`, so only the caps bound the final total. All five are also per-request on the `fetch_url` MCP tool (`limit`, `depth`, `sort`, `max_comments`, `max_top_level`); a positive value overrides the env default.\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"100%\"\u003e\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Travel%20and%20places/Building%20Construction.png\" width=\"26\" height=\"26\" /\u003e Architecture\n\nTwo ports answer two questions. The `Searcher` port answers *query → URLs*; the `Engine` port answers *URL → content*. MCP tools and REST handlers compose them; transports stay thin.\n\n```mermaid\n%%{init: {\"theme\":\"base\",\"themeVariables\":{\"background\":\"transparent\",\"mainBkg\":\"#161b22\",\"primaryColor\":\"#161b22\",\"primaryTextColor\":\"#e6edf3\",\"primaryBorderColor\":\"#FF4500\",\"lineColor\":\"#8b949e\",\"secondaryColor\":\"#161b22\",\"tertiaryColor\":\"#161b22\"},\"flowchart\":{\"curve\":\"basis\",\"htmlLabels\":false}}}%%\nflowchart TB\n  crawl[\"POST /crawl\"] e1@--\u003e owt[\"Open WebUI\u003cbr/\u003etransport\"]\n  search[\"POST /search\"] e2@--\u003e sat[\"SearchAPI\u003cbr/\u003etransport\"]\n  mcpStdio[\"MCP stdio\"] e3@--\u003e mcp[\"MCP server\"]\n  mcpHTTP[\"MCP HTTP /mcp\"] e4@--\u003e mcp\n\n  owt e5@--\u003e reg[\"Engine Registry\"]\n  mcp -- crawl tools --\u003e reg\n  sat e6@--\u003e searcher[\"Searcher\u003cbr/\u003e(SearXNG)\"]\n  mcp -- search tool --\u003e searcher\n\n  reg e7@--\u003e reddit[\"Reddit engine\u003cbr/\u003e(TOON)\"]\n  reg e12@--\u003e hn[\"Hacker News engine\u003cbr/\u003e(TOON)\"]\n  reg e8@--\u003e generic[\"Generic fallback\u003cbr/\u003e(markdown)\"]\n  reddit e9@--\u003e c4[\"crawl4ai upstream\u003cbr/\u003e(headless browser)\"]\n  generic e10@--\u003e c4\n  hn e13@--\u003e algolia[\"Algolia HN API\u003cbr/\u003e(hn.algolia.com)\"]\n  searcher e11@--\u003e sx[\"SearXNG upstream\u003cbr/\u003e(Google / Bing / DDG)\"]\n\n  classDef box fill:#161b22,stroke:#30363d,stroke-width:1px,color:#e6edf3;\n  classDef accent fill:#0d1117,stroke:#FF4500,stroke-width:2px,color:#ffd9b3;\n  classDef animate stroke:#FF4500,stroke-width:2px,stroke-dasharray:10 6,stroke-dashoffset:900,animation:dash 14s linear infinite;\n  class crawl,search,mcpStdio,mcpHTTP,owt,sat box;\n  class mcp,reg,searcher,reddit,hn,generic,c4,sx,algolia accent;\n  class e1,e2,e3,e4,e5,e6,e7,e8,e9,e10,e11,e12,e13 animate;\n```\n\n### \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Shield.png\" width=\"22\" height=\"22\" /\u003e Reddit anti-bot handling\n\nReddit's edge 403-blocks non-browser HTTP clients (it fingerprints the TLS/JA3 handshake), so the Reddit engine never calls Reddit directly. It drives crawl4ai's headless Chromium to a `www.reddit.com` page (which clears the bot wall), then runs a **same-origin `fetch()`** of the `.json` and `/api/morechildren` endpoints from inside that page. No Reddit auth, cookies, or API key. A per-thread crawl4ai `session_id` is reused across expansion rounds to keep one warmed context.\n\n\u003e Sustained scraping can raise your source IP's risk score. If fetches start returning the block page, slow down, keep `expand` modest, or route crawl4ai through a residential proxy.\n\n### \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Activities/Puzzle%20Piece.png\" width=\"22\" height=\"22\" /\u003e Extending it\n\nNew engines (Hacker News, Stack Overflow, …), searchers (Brave, Tavily, …), MCP tools, and transports each plug into one small port without touching the rest. See **[AGENTS.md → Adding things](AGENTS.md#adding-things)** for the architecture and a step-by-step.\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"100%\"\u003e\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Hammer%20and%20Wrench.png\" width=\"26\" height=\"26\" /\u003e Development\n\n```bash\ngit clone https://github.com/kinorai/omnifeed.git \u0026\u0026 cd omnifeed\nmake check        # vet + lint + test — hermetic, no upstreams or token needed\ndocker compose up # run the full stack locally (tokenless: ports 8080 / 8081 / 9090)\n```\n\nSee [CONTRIBUTING.md](CONTRIBUTING.md) for the full workflow, [SECURITY.md](SECURITY.md) for vulnerability reporting, and [AGENTS.md](AGENTS.md) if you're a coding agent working in this repo.\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"100%\"\u003e\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Travel%20and%20places/Star.png\" width=\"26\" height=\"26\" /\u003e Star history\n\n\u003ca href=\"https://star-history.com/#kinorai/omnifeed\u0026Date\"\u003e\n \u003cpicture\u003e\n   \u003csource media=\"(prefers-color-scheme: dark)\" srcset=\"https://api.star-history.com/svg?repos=kinorai/omnifeed\u0026type=Date\u0026theme=dark\" /\u003e\n   \u003csource media=\"(prefers-color-scheme: light)\" srcset=\"https://api.star-history.com/svg?repos=kinorai/omnifeed\u0026type=Date\" /\u003e\n   \u003cimg alt=\"Star History Chart\" src=\"https://api.star-history.com/svg?repos=kinorai/omnifeed\u0026type=Date\" width=\"70%\" /\u003e\n \u003c/picture\u003e\n\u003c/a\u003e\n\n\u003cimg src=\"https://user-images.githubusercontent.com/74038190/212284100-561aa473-3905-4a80-b561-0d28506553ee.gif\" width=\"100%\"\u003e\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Hand%20gestures/Handshake.png\" width=\"26\" height=\"26\" /\u003e Contributing\n\n\u003cdiv align=\"center\"\u003e\n\n**Star it if it's useful — it helps other AI builders find omnifeed.**\n\n[![Star](https://img.shields.io/badge/⭐_Star_omnifeed-FF4500?style=for-the-badge)](https://github.com/kinorai/omnifeed)\n[![Open an issue](https://img.shields.io/badge/🐛_Open_an_Issue-161b22?style=for-the-badge)](https://github.com/kinorai/omnifeed/issues/new)\n[![Submit a PR](https://img.shields.io/badge/🔧_Submit_a_PR-7C3AED?style=for-the-badge)](https://github.com/kinorai/omnifeed/pulls)\n\n\u003c/div\u003e\n\nNew engines, searchers, MCP tools, and transports are all welcome — start with [AGENTS.md](AGENTS.md#adding-things) and [CONTRIBUTING.md](CONTRIBUTING.md).\n\n## \u003cimg src=\"https://raw.githubusercontent.com/Tarikul-Islam-Anik/Animated-Fluent-Emojis/master/Emojis/Objects/Page%20Facing%20Up.png\" width=\"26\" height=\"26\" /\u003e License\n\n[MIT](LICENSE) © kinorai\n\n\u003cimg src=\"https://capsule-render.vercel.app/api?type=waving\u0026color=0:7C3AED,100:FF4500\u0026height=120\u0026section=footer\" width=\"100%\"/\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkinorai%2Fomnifeed","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkinorai%2Fomnifeed","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkinorai%2Fomnifeed/lists"}