{"id":51958169,"url":"https://github.com/KnockOutEZ/wigolo","last_synced_at":"2026-08-04T07:00:35.183Z","repository":{"id":350895660,"uuid":"1208642537","full_name":"KnockOutEZ/wigolo","owner":"KnockOutEZ","description":"The go-to web for your AI coding agent — local-first search, fetch, crawl \u0026 research over MCP. No API keys, no cloud, $0/query. Public beta.","archived":false,"fork":false,"pushed_at":"2026-07-28T10:39:11.000Z","size":32128,"stargazers_count":3792,"open_issues_count":27,"forks_count":251,"subscribers_count":15,"default_branch":"main","last_synced_at":"2026-07-28T12:10:05.417Z","etag":null,"topics":["agent","ai","ai-agent","claude","cli","developer-tools","local-first","mcp","mcp-server","metasearch","model-context-protocol","nodejs","privacy","rag","search","search-engine","typescript","web-crawler","web-scraping","web-search"],"latest_commit_sha":null,"homepage":"https://knockoutez.github.io/wigolo/","language":"TypeScript","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/KnockOutEZ.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":".github/FUNDING.yml","license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null},"funding":{"buy_me_a_coffee":"knockoutez"}},"created_at":"2026-04-12T15:04:11.000Z","updated_at":"2026-07-28T11:50:20.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/KnockOutEZ/wigolo","commit_stats":null,"previous_names":["knockoutez/wigolo"],"tags_count":61,"template":false,"template_full_name":null,"purl":"pkg:github/KnockOutEZ/wigolo","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/KnockOutEZ%2Fwigolo","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/KnockOutEZ%2Fwigolo/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/KnockOutEZ%2Fwigolo/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/KnockOutEZ%2Fwigolo/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/KnockOutEZ","download_url":"https://codeload.github.com/KnockOutEZ/wigolo/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/KnockOutEZ%2Fwigolo/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36265484,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-08-04T02:00:06.901Z","response_time":57,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agent","ai","ai-agent","claude","cli","developer-tools","local-first","mcp","mcp-server","metasearch","model-context-protocol","nodejs","privacy","rag","search","search-engine","typescript","web-crawler","web-scraping","web-search"],"created_at":"2026-07-29T13:00:23.717Z","updated_at":"2026-08-04T07:00:35.155Z","avatar_url":"https://github.com/KnockOutEZ.png","language":"TypeScript","funding_links":["https://buymeacoffee.com/knockoutez"],"categories":["5. Retrieval-Augmented Generation (RAG) \u0026 Knowledge","القائمة الكاملة Top 200","Repos"],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n\u003cimg alt=\"wigolo — the go-to web for your agent\" src=\"assets/brand/wigolo-banner.png\" width=\"640\"\u003e\n\nLocal-first web intelligence for AI agents — **no keys, no cloud, no metered bill.**\n\n\u003csub\u003eworks with\u0026nbsp;\u0026nbsp;**Claude Code · Cursor · Codex · Gemini CLI · OpenCode · VS Code · Windsurf · Zed · Antigravity**\u003c/sub\u003e\n\u003cbr\u003e\n\u003csub\u003eand beyond\u0026nbsp;\u0026nbsp;**LangChain · CrewAI · LlamaIndex · Vercel AI SDK · n8n \u0026 self-hosted agents · any MCP client · plain REST**\u003c/sub\u003e\n\n[![npm](https://img.shields.io/npm/v/wigolo?color=cb3837\u0026logo=npm)](https://www.npmjs.com/package/wigolo)\n[![npm downloads](https://img.shields.io/npm/dm/wigolo?color=cb3837\u0026logo=npm\u0026label=downloads)](https://www.npmjs.com/package/wigolo)\n[![GitHub stars](https://img.shields.io/github/stars/KnockOutEZ/wigolo?style=flat\u0026logo=github\u0026color=e3b341)](https://github.com/KnockOutEZ/wigolo/stargazers)\n[![CI](https://img.shields.io/github/actions/workflow/status/KnockOutEZ/wigolo/ci.yml?branch=main\u0026logo=github\u0026label=CI)](https://github.com/KnockOutEZ/wigolo/actions/workflows/ci.yml)\n[![node](https://img.shields.io/badge/node-%E2%89%A520-339933?logo=node.js\u0026logoColor=white)](https://nodejs.org)\n[![MCP](https://img.shields.io/badge/MCP-server-7c3aed)](https://modelcontextprotocol.io)\n[![license](https://img.shields.io/badge/license-AGPL--3.0-2563eb)](#license)\n[![status](https://img.shields.io/badge/status-public%20beta-b7791f)](#beta--feedback)\n[![follow on X](https://img.shields.io/badge/follow-%40yourtowhid-000000?logo=x\u0026logoColor=white)](https://x.com/yourtowhid)\n\n\u003ca href=\"https://trendshift.io/repositories/79424?utm_source=repository-badge\u0026utm_medium=badge\u0026utm_campaign=badge-repository-79424\" target=\"_blank\"\u003e\u003cimg src=\"https://trendshift.io/api/badge/repositories/79424\" alt=\"wigolo on Trendshift\" width=\"250\" height=\"55\"/\u003e\u003c/a\u003e\n\u003ca href=\"https://trendshift.io/repositories/79424?utm_source=trendshift-badge\u0026utm_medium=badge\u0026utm_campaign=badge-trendshift-79424\" target=\"_blank\" rel=\"noopener noreferrer\"\u003e\u003cimg src=\"https://trendshift.io/api/badge/trendshift/repositories/79424/daily?language=TypeScript\" alt=\"KnockOutEZ%2Fwigolo | Trendshift\" width=\"250\" height=\"55\"/\u003e\u003c/a\u003e\n\n[Quickstart](#quickstart) · [Tools](#tools) · [Why wigolo](#why-its-different) · [Benchmark](#benchmark) · [Docs](docs/README.md) · [Examples](examples/README.md) · [Feedback](#beta--feedback) · [FAQ](#faq)\n\nNew features and updates ship steadily. Follow \u003ca href=\"https://x.com/yourtowhid\"\u003e\u003cb\u003e@yourtowhid on X\u003c/b\u003e\u003c/a\u003e for all of it and new ways to use wigolo, and reach out there for collaborations or feedback · also on \u003ca href=\"https://www.linkedin.com/in/yourtowhid/\"\u003eLinkedIn\u003c/a\u003e\n\n\u003c/div\u003e\n\n---\n\nwigolo gives an AI agent one surface for everything web-related: **search, fetch, crawl, extract, cache, find-similar, research,** and autonomous gather loops. It runs wherever your agent runs — as an MCP server next to your coding agent, as a REST/MCP endpoint on the box where your self-hosted agents live, or embedded through an SDK inside your own app. The core tools need no API keys, nothing it touches leaves `~/.wigolo/`, and no bill grows with how much your agent thinks.\n\n\u003cdiv align=\"center\"\u003e\n\n\u003cimg alt=\"wigolo demo — Claude Code answering a live web question through wigolo, no API keys\" src=\"assets/wigolo-demo.gif\" width=\"800\"\u003e\n\n\u003c/div\u003e\n\n## Quickstart\n\n```bash\nnpx wigolo init                              # set up the local engine — any system\nnpx wigolo init --agents=claude-code,cursor  # …or set up + wire your day-to-day agents in one command\n```\n\nRequires **Node ≥ 20** and ~1.5 GB of free disk on macOS, Linux, or Windows. Bare `init` sets up the local engine: it downloads the browser engine and on-device models, runs a health check, and reports each component. Adding `--agents` wires the named agents in the same run, so a coding agent you use daily is ready in one command.\n\n- **Supported agents** — `--agents` takes any of `claude-code` · `cursor` · `codex` · `gemini-cli` · `opencode` · `vscode` · `windsurf` · `zed` · `antigravity` (comma-separated); wigolo writes the MCP config and, where supported, instructions for each.\n- **Any other setup** — any MCP client, agent framework, or self-hosted agent registers `npx -y wigolo` in its own MCP config. The [installation guide](docs/installation.md) has the exact config block for every client, plus Docker, Homebrew, and single-file-binary channels.\n- **More on the way** — the supported list keeps growing, and a PR to add your agent is welcome; see [CONTRIBUTING.md](CONTRIBUTING.md).\n- **Interactive setup** — `--interactive` is a plain-text flow; `--wizard` is the full terminal TUI.\n- **Defer downloads** — `--no-warmup` waits until first use. A failed component download never fails setup; init reports what's not ready with the exact fix and still completes.\n\n`init` is unattended by default, so it's safe in scripts and CI, and any setup problem surfaces right here in the per-component report, before your agent's first call. **Search, fetch, crawl, extract, cache, and find-similar work with no API key.** Check it's healthy anytime:\n\n```bash\nnpx wigolo doctor\n```\n\nTo remove everything cleanly, run `npx wigolo config --uninstall --yes`. You can also paste the [installation guide](docs/installation.md) into any AI assistant and let it do the setup; it's written to be self-contained.\n\n### Recommended — a free key for `research` \u0026 `agent`\n\nSearch, fetch, crawl, extract, cache, and find-similar are **fully keyless**. `research`, `agent`, and `search format=answer` use an LLM to write the synthesized, cited answer. Without one they hand back a raw brief and evidence for your agent to assemble. A free Gemini key turns that into a finished answer:\n\n```bash\nexport WIGOLO_LLM_PROVIDER=gemini\nexport GEMINI_API_KEY=\u003cfree-key\u003e      # grab one at aistudio.google.com/apikey — the free tier is plenty\n```\n\nAny provider works (`anthropic` · `openai` · `groq`), or stay fully local and keyless with `WIGOLO_LLM_PROVIDER=ollama` (or any OpenAI-compatible URL). Set it in your shell or your agent's MCP `env` block. Providers, models, and the keyless local-model ladder are in the [configuration guide](docs/configuration.md).\n\n## What your agent gets back\n\nEvery search result is evidence the agent can act on. It carries a verbatim excerpt pinned to its exact position in the source, a citation ID the agent can quote, and a score it can inspect (abridged real shape):\n\n```jsonc\n{\n  \"results\": [{\n    \"title\": \"Logical replication - PostgreSQL docs\",\n    \"url\": \"https://www.postgresql.org/docs/current/logical-replication.html\",\n    \"excerpt\": \"Logical replication is a method of replicating data objects…\",\n    \"citation_id\": \"src-1\",\n    \"source_span\": { \"start\": 1042, \"end\": 1305 },          // byte-exact provenance\n    \"evidence_score\": { \"final\": 0.86, \"semantic\": 0.91, \"lexical\": 0.78, \"engine_consensus\": 3 }\n  }],\n  \"citations\": [{ \"id\": \"src-1\", \"url\": \"…\" }],\n  \"freshness_signal\": { \"published\": \"2026-05-12\", \"confidence\": \"high\" }\n}\n```\n\nWeak results get flagged as junk by wigolo's own scorer. Failed engines are reported and stale cache is labeled, so the agent always knows what it's standing on. Full response contracts per tool are in the [tools reference](docs/tools.md).\n\n## Tools\n\n| Tool | What it does |\n|------|--------------|\n| 🔎 `search` | Multi-engine web search (18 direct adapters) with rank fusion, ML reranking, and an explainable per-result score. Pass a query **array** for parallel breadth. Scope by domain and time range, match an exact phrase, or return image results. |\n| 📄 `fetch` | Load one URL through a tiered router that auto-escalates from plain HTTP to a headless browser engine on anti-bot challenges or SPA shells. Clean markdown + metadata + links. Handles PDFs, a single-heading `section`, authenticated sessions, and page actions (click / type / scroll / screenshot). |\n| 🕸️ `crawl` | Multi-page crawl — BFS, DFS, sitemap, or map-only. Per-domain rate limits, robots.txt respect, boilerplate dedup. |\n| 🧩 `extract` | Structured data from a page: tables, metadata, JSON-LD, brand identity, named schemas (Article / Recipe / Product / …), or any custom JSON Schema. |\n| 💾 `cache` | Query everything already seen — keyword or hybrid semantic. Plus stats, clear, and change detection. |\n| 🧲 `find_similar` | Pages similar to a URL or a concept, via 3-way fusion of keyword + semantic + live web. |\n| 🧠 `research` | Decompose a question → fan out sub-queries → fetch sources → synthesize a cited report (or a structured brief the host LLM writes from). |\n| 🤖 `agent` | Autonomous gather loop: plan → search → fetch → extract → synthesize, with a step log, time budget, and optional output schema. |\n| 🔁 `diff` + ⏱️ `watch` | See exactly what changed on a page since last visit; re-check on demand and deliver changes to a webhook. |\n\nEvery tool also runs from the terminal (`wigolo search \"…\" --json`), from an interactive shell with NDJSON piping (`wigolo shell`), over REST, and through the SDKs — [CLI reference](docs/cli.md). Per-tool guides with the full parameter set are in [docs/tools.md](docs/tools.md); runnable examples are in [examples/](examples/README.md).\n\n## Why it's different\n\nwigolo isn't a free stand-in for the paid tools — it's built to match them. It's a focused web layer for your agents: an MCP and REST surface they call directly, with the search and extraction quality the paid services charge for. What separates it:\n\n- **Built for agents.** One MCP call fans out many queries across many engines in parallel, which a serial host tool-loop can't replicate. Every result carries transparent per-result scoring, and output is budget-aware.\n- **Honest output.** Stale cache, failed fetches, degraded backends, and truncation are surfaced in the result. When a bot-protected page can't be read, you get a labeled `blocked_by_challenge` failure, not a challenge shell returned as content.\n- **$0 per query, free to re-query.** Default search talks to public engines through direct adapters; the reranker and embeddings run on-device. Every response is cached, so asking again is instant and costs nothing.\n- **Private by default.** Cache, embeddings, models, and config live under `~/.wigolo/`. Nothing reaches a third party unless you explicitly opt into an LLM for synthesis.\n\nHere's what one real result looks like, dissected. It includes the failed engine and the weak result, because those are part of the answer too:\n\n\u003cdiv align=\"center\"\u003e\n\n\u003cpicture\u003e\n\u003csource media=\"(prefers-color-scheme: dark)\" srcset=\"assets/promo/anatomy-dark.svg\"\u003e\n\u003cimg alt=\"Anatomy of a wigolo result: explainable score decomposition, live engine telemetry, surfaced degradation, self-flagged junk — one real query, captured live\" src=\"assets/promo/anatomy.svg\" width=\"880\"\u003e\n\u003c/picture\u003e\n\n\u003c/div\u003e\n\n## Benchmark\n\n\u003e **All four tools converged on the same core answer, and only one of them handed back verbatim, byte-pinned evidence while doing it.**\n\nOne cold query ran live inside a single **Claude Fable 5** session, fanned out to four web tools on equal footing (built-in **WebSearch**, **wigolo**, **Tavily**, **Exa**), and was judged by the agent on the evidence alone. All four converged on the same answer and the same top source, so the parity is demonstrated on-screen. wigolo alone returned verbatim excerpts pinned to byte-offset source spans, an explainable score decomposition, and live per-engine telemetry, and its own scorer flagged two weak results as junk. The cloud tools earn their place too: Exa rendered the official docs' comparison matrix in full. Run your own query and you'll see the same shape.\n\n\u003cdiv align=\"center\"\u003e\n\n\u003cimg alt=\"wigolo vs built-in WebSearch, Tavily, and Exa on one real query, driven by Claude Fable 5\" src=\"assets/wigolo-vs.gif\" width=\"900\"\u003e\n\n\u003c/div\u003e\n\n### How it compares\n\n| | wigolo | Firecrawl | Exa | Tavily |\n|---|:---:|:---:|:---:|:---:|\n| Multi-engine web search | ✅ | ✅ | ✅ | ✅ |\n| Fetch \u0026 structured extraction | ✅ | ✅ | ✅ | ✅ |\n| Whole-site crawl \u0026 map | ✅ | ✅ | — | ✅ |\n| Verbatim excerpts pinned to byte-offset source spans | ✅ | — | — | — |\n| Explainable per-result score decomposition | ✅ | — | — | — |\n| Persistent local memory — re-query instantly, offline | ✅ | — | — | — |\n| Query data stays on your machine | ✅ | — | — | — |\n| API key / account | none | required | required | required |\n| Cost per query | $0 | metered | metered | metered |\n\n\u003csub\u003eFeature standing as of July 2026 — check each vendor's docs for current state.\u003c/sub\u003e\n\nThat last row compounds, because agents ask in bursts:\n\n\u003cdiv align=\"center\"\u003e\n\n\u003cpicture\u003e\n\u003csource media=\"(prefers-color-scheme: dark)\" srcset=\"assets/promo/meter-dark.svg\"\u003e\n\u003cimg alt=\"The meter: a metered cloud API's cost climbs with every query while wigolo stays flat at zero dollars — illustrative pricing\" src=\"assets/promo/meter.svg\" width=\"880\"\u003e\n\u003c/picture\u003e\n\n\u003c/div\u003e\n\n## Beyond your editor\n\nThe same ten tools serve every kind of agent, over whichever surface fits: MCP for coding agents, REST for everything else, SDKs to embed, and framework wrappers to drop in.\n\n### REST API — `wigolo serve`\n\nOne process exposes a plain-JSON REST API next to the MCP transport. No MCP client needed, just curl:\n\n```bash\nwigolo serve                          # 127.0.0.1:3333 — loopback is open; off-loopback requires a token\n\ncurl -sX POST http://127.0.0.1:3333/v1/search \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"query\":\"local-first software\",\"max_results\":5}'\n```\n\n`POST /v1/{tool}` covers all ten tools, `GET /openapi.json` is the OpenAPI 3.1 contract, and `/mcp` + `/sse` serve remote MCP clients from the same port. Bind past loopback and a bearer token is required, so the server fails closed by default. Point n8n, a Hermes-style assistant, or any self-hosted agent at it. → [REST API](docs/rest-api.md)\n\n### SDKs — TypeScript \u0026 Python\n\nThin, typed clients with an embedded local mode that finds or starts the daemon for you. No separate `serve` step.\n\n**TypeScript** — `npm install wigolo-sdk` (zero-dep; Node / Bun / Deno / edge):\n\n```ts\nimport { createLocalClient } from 'wigolo-sdk/local';\n\nconst { client, close } = await createLocalClient();   // reuse a running daemon, or spawn one\nconst res = await client.search({ query: 'local-first web search', max_results: 5 });\nconsole.log(res.results.map((r) =\u003e r.title));\nawait close();                                          // stops the daemon only if this call spawned it\n```\n\n**Python** — `pip install wigolo` (standard library only; sync + async):\n\n```python\nfrom wigolo import local_client\n\nwith local_client() as client:                          # reuse a healthy daemon, or spawn one\n    res = client.search(query=\"local-first web search\", max_results=5)\n    for r in res[\"results\"]:\n        print(r[\"title\"], r[\"url\"])\n```\n\n→ [SDKs \u0026 embedded mode](docs/sdks.md)\n\n### Framework integrations\n\nDrop wigolo's tools into the framework you already use. You get the full ten-tool surface, including the cache / find_similar / research / agent that most framework web-tools don't ship:\n\n| Framework | Package | What you get |\n|-----------|---------|--------------|\n| **LangChain** | `wigolo-langchain` | each tool as a `BaseTool`, plus a `BaseRetriever` over search / find_similar for RAG |\n| **CrewAI** | `wigolo-crewai` | `wigolo_tools()` → hand the set to any crew |\n| **LlamaIndex** | `wigolo-llamaindex` | a `BaseReader` that loads fetched / crawled / searched pages as documents |\n| **Vercel AI SDK** | `wigolo-vercel-ai-sdk` | tool factories for `generateText` / `streamText`, edge-friendly |\n\n→ [Framework integrations](docs/sdks.md)\n\n### Docker\n\n```bash\n# stdio MCP — wire it into any MCP client as command: docker\ndocker run -i --rm -v wigolo-data:/data ghcr.io/knockoutez/wigolo\n\n# HTTP server for remote / multi-client use\ndocker run -p 3333:3333 -v wigolo-data:/data \\\n  -e WIGOLO_API_TOKEN=a-long-random-secret \\\n  ghcr.io/knockoutez/wigolo serve --host 0.0.0.0\n```\n\nThe slim image lazy-loads models into the volume; `:full` preinstalls the browser engine. Also on Docker Hub as `towhid69420/wigolo`. → [installation \u0026 all channels](docs/installation.md)\n\n### Agent skills\n\nAn 11-pack skill catalog teaches your coding agent to drive each tool well. It's installed by `init` and managed with `wigolo skills add|list|remove`. → [skills](docs/skills.md)\n\nOne note for self-hosters: some challenge-protected sites score IP reputation, so a datacenter IP won't clear walls a home connection would. wigolo labels those failures, and the [self-hosting guide](docs/self-hosting.md) covers the opt-in proxy answer.\n\n## Star history\n\n\u003cdiv align=\"center\"\u003e\n\n\u003ca href=\"https://www.star-history.com/#KnockOutEZ/wigolo\u0026Date\"\u003e\n\u003cpicture\u003e\n\u003csource media=\"(prefers-color-scheme: dark)\" srcset=\"https://raw.githubusercontent.com/KnockOutEZ/wigolo/star-chart/star-history-dark.svg\"\u003e\n\u003cimg alt=\"wigolo GitHub stars over time\" src=\"https://raw.githubusercontent.com/KnockOutEZ/wigolo/star-chart/star-history.svg\" width=\"880\"\u003e\n\u003c/picture\u003e\n\u003c/a\u003e\n\n\u003csub\u003eRefreshed daily from the GitHub API. \u003ca href=\"https://github.com/KnockOutEZ/wigolo\"\u003eAdd a ⭐\u003c/a\u003e if wigolo is useful to you.\u003c/sub\u003e\n\n\u003c/div\u003e\n\n## Architecture\n\nA single Node process speaks MCP (JSON-RPC over stdio). Everything heavy is local and lazy-loaded, so a zero-key install pays nothing for the parts it isn't using.\n\n```mermaid\nflowchart TD\n    A[\"🤖 AI agent\u003cbr/\u003eany MCP client · REST · SDK\"]\n    A --\u003e|MCP over stdio| B[\"\u003cb\u003ewigolo\u003c/b\u003e\u003cbr/\u003e10 tools · dynamic instructions\u003cbr/\u003ein-process browser pool + cache + models\"]\n\n    B --\u003e C{\"Tool layer\"}\n    C --\u003e T1[\"search · fetch · crawl · extract\"]\n    C --\u003e T2[\"cache · find_similar · research · agent\"]\n\n    T1 --\u003e F[\"⚙️ Fetch router\u003cbr/\u003etiered escalation, learned per domain\"]\n    T1 --\u003e S[\"⚙️ Search\u003cbr/\u003e18 engines → rank fusion → ML rerank\u003cbr/\u003e\u003ci\u003eexplainable evidence score\u003c/i\u003e\"]\n    T2 --\u003e DB[(\"🗄️ Local cache\u003cbr/\u003ekeyword + vector index\")]\n    T2 --\u003e ML[\"🧠 On-device ML\u003cbr/\u003eembeddings + reranker\"]\n\n    F -.-\u003e|optional| LLM[\"☁️ LLM\u003cbr/\u003esynthesis only · opt-in\"]\n    S -.-\u003e|optional| SX[\"🔀 Aggregator backend\u003cbr/\u003eopt-in legacy / hybrid\"]\n\n    F --\u003e WEB[\"🌍 Public web\"]\n    S --\u003e WEB\n\n    style B fill:#7c3aed,stroke:#5b21b6,color:#fff\n    style WEB fill:#0ea5e9,stroke:#0369a1,color:#fff\n    style DB fill:#1e293b,stroke:#334155,color:#fff\n    style LLM stroke-dasharray: 5 5\n    style SX stroke-dasharray: 5 5\n```\n\n- **Code beats model.** Deterministic work stays off the LLM: canonicalization, rank fusion, dedup, and schema matching. The model is reserved for judgment, opt-in, and capped per request. LLM-filled fields are checked against the source and nulled if absent.\n- **Signal-driven routing.** The fetch ladder escalates to a real browser on observable signals, not domain guesses: SPA markers, challenge bodies, thin content. It learns per domain, unlearns when a site stops needing it, and `wigolo tune list` shows you exactly what it learned.\n- **Reads pages the way a browser does.** Tiered fetching waits out interstitial challenges and reuses clearances per domain, politely: robots.txt respected, per-domain rate limits, research-grade volumes. When a wall stays up, the failure is labeled and reported.\n\n## Configuration\n\nA clean install works out of the box. Three settings raise output quality:\n\n```bash\n# 1. Synthesis — the biggest lever (research / agent / search-answer write real prose)\nexport WIGOLO_LLM_PROVIDER=gemini                   # or anthropic / openai / groq / ollama (keyless)\nexport GEMINI_API_KEY=\u003cyour-key\u003e\n\n# 2. Wider retrieval funnel\nexport WIGOLO_SEARCH=hybrid                         # core engines + aggregator fallback\nexport WIGOLO_GITHUB_TOKEN=...                      # GitHub code search 10 → 30 req/min\n\n# 3. Land more fetches, stay warm\nexport WIGOLO_TLS_TIER=auto                         # per-domain learned fetch hardening\nexport WIGOLO_EAGER_WARMUP=1                        # pay the ~1s model load up front\n```\n\n**Per-call habits that pay off:** query **arrays** (`[\"a\",\"b\",\"c\"]`) for parallel breadth · `search_depth: \"deep\"` for queries that matter · `include_domains` as a hard filter for docs lookups. The full reference covers every environment variable, config-file key, search backend, cache TTL, and serve limit; it's in the [configuration guide](docs/configuration.md).\n\n## Docs \u0026 examples\n\n**[docs/](docs/README.md)** — the complete manual:\n[getting started](docs/getting-started.md) · [installation \u0026 channels](docs/installation.md) · [configuration](docs/configuration.md) · [tools reference](docs/tools.md) · [CLI \u0026 shell](docs/cli.md) · [REST API](docs/rest-api.md) · [SDKs \u0026 integrations](docs/sdks.md) · [self-hosting](docs/self-hosting.md) · [agent skills](docs/skills.md) · [plugins](docs/plugins.md) · [troubleshooting \u0026 FAQ](docs/troubleshooting.md) · [privacy \u0026 security](docs/privacy-security.md)\n\n**[examples/](examples/README.md)** — runnable, each with a README (and most with a terminal recording): one-shot CLI, NDJSON shell pipelines, REST via curl, TypeScript \u0026 Python SDKs, Vercel AI SDK tools, pointing self-hosted n8n at a remote wigolo, watch-with-webhook, and writing your own search-engine plugin. The docs are also rendered on the site at **[knockoutez.github.io/wigolo/docs](https://knockoutez.github.io/wigolo/docs/)**.\n\n## Beta \u0026 feedback\n\nwigolo is in **public beta**. Everything documented here works and is held to a 7,600-test suite; it's stable, and beta is about the polish bar. It stays beta until enough people have used it, kicked it, and starred it that calling it v1 means something. Your feedback shapes what comes next, and every report is read, usually the same day:\n\n- 🐛 **[Report a bug](https://github.com/KnockOutEZ/wigolo/issues/new?template=bug_report.yml)** — broke, misbehaved, surprised you\n- 💡 **[Request a feature](https://github.com/KnockOutEZ/wigolo/issues/new?template=feature_request.yml)** — something it should do\n- 💬 **[Ask anything](https://github.com/KnockOutEZ/wigolo/discussions)** — questions, setups, show \u0026 tell\n\nIf wigolo earns a place in your setup, three things keep it going: a ⭐ **star** (it's how open source gets found), a **[☕ coffee](https://buymeacoffee.com/knockoutez)** (there's no paid tier and never will be), or **[an email](mailto:ktowhid20@gmail.com)** that goes straight to the one developer who wrote the code.\n\n## Troubleshooting\n\n`wigolo doctor` names any broken component and the exact env var or command that fixes it; `wigolo doctor --fix` repairs the common cases, and `wigolo verify` health-checks every component. A component failing during `init` doesn't break wigolo: `init` still exits 0, and core search / fetch / crawl / extract / cache work with no models and no browser. Quick hits:\n\n- **Slow or failed downloads** — re-run `wigolo warmup --all` (or `--browser` / `--embeddings` / `--reranker`); they resume and retry.\n- **Browser won't launch on Linux** — `wigolo warmup --browser` installs the OS libraries (or prints the exact command).\n- **Native build error / unusual Node** — use an LTS: **Node 20, 22, or 24**.\n- **Behind a proxy** — `USE_PROXY=true` + `PROXY_URL`; add `NODE_EXTRA_CA_CERTS` for TLS-inspecting proxies.\n\nThe full guide covers per-symptom fixes, a \"what still works when X fails\" map, platform notes (incl. linux-arm64), and offline installs: **[docs/troubleshooting.md](docs/troubleshooting.md)**.\n\n## FAQ\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eFree? What's the catch?\u003c/b\u003e\u003c/summary\u003e\n\nNo catch by design. The expensive parts (ranking, embeddings, the browser engine) run on *your* hardware, so there's no per-query cost to recover and no reason for a meter. It's sustained by donations, and the AGPL license legally prevents a switch into a closed hosted product.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eIs the quality really on par with the paid services?\u003c/b\u003e\u003c/summary\u003e\n\nThe benchmark section above is a live 4-way run you can reproduce: everyday agent queries land at parity, the paid tools still win some deep-extraction edge cases, and crawling is where wigolo is strongest. Every result shows its scoring, so you don't have to take anyone's word for it.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eWon't public search engines block or rot?\u003c/b\u003e\u003c/summary\u003e\n\nIt's engineered for exactly that: 18 engines fused with rank fusion (any one failing barely moves results), a tiered fetch ladder with per-domain learning, and an optional aggregator fallback. Degraded backends are *reported in the output*, and the local cache means everything already seen keeps working regardless.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eIs this kind of scraping OK?\u003c/b\u003e\u003c/summary\u003e\n\nwigolo reads the public web the way a browser does: robots.txt respected by default, per-domain rate limits, and research-grade volumes for one agent on one machine. It sits deliberately at the polite end of the spectrum.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eAGPL — can I use this at work?\u003c/b\u003e\u003c/summary\u003e\n\nYes, freely, company-wide. The license only bites if you *modify wigolo and run it as a network service*, in which case you must publish those modifications; using it as a local dev tool carries zero obligation. For commercial-licensing questions, reach out.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eWhy 1.5 GB of disk?\u003c/b\u003e\u003c/summary\u003e\n\nThat's the on-device brain: a full browser engine plus the ranking and embedding models the cloud services run on their side and bill you for. Once it's on disk, every query uses it for free.\n\n\u003c/details\u003e\n\n## Available on\n\n- **npm** — [`wigolo`](https://www.npmjs.com/package/wigolo) *(primary channel — the Quickstart above)*\n- **PyPI** — [`wigolo`](https://pypi.org/project/wigolo/) *(Python SDK)*\n- **Docker** — [`ghcr.io/knockoutez/wigolo`](https://github.com/KnockOutEZ/wigolo/pkgs/container/wigolo) · [`towhid69420/wigolo`](https://hub.docker.com/r/towhid69420/wigolo)\n- **Official MCP Registry** — `io.github.KnockOutEZ/wigolo`\n- **Directories** — [Glama](https://glama.ai/mcp/servers/KnockOutEZ/wigolo) · [Smithery](https://smithery.ai/server/ktowhid20/wigolo) · [mcp.so](https://mcp.so/server/wigolo/KnockOutEZ) · [LobeHub](https://lobehub.com/mcp/knockoutez-wigolo)\n\nHomebrew, `curl | sh`, and the single-file binary are covered in the [installation guide](docs/installation.md). Use one channel per machine; they all share `~/.wigolo`.\n\n## Contributing\n\nBug reports, feature requests, and PRs are all welcome; see **[CONTRIBUTING.md](CONTRIBUTING.md)**. Keep tool handlers thin, add tests, and run the suite before opening a PR. The friendliest entry point is the plugin system for custom search engines and extractors: [add a search engine in ~100 lines](docs/plugins.md), with a template in [`examples/plugin-search-engine`](examples/plugin-search-engine).\n\n## License\n\n**[GNU AGPL-3.0-only](LICENSE).** Free to use, modify, and self-host, including inside a company. The one obligation: if you run a **modified** version as a network service, you must publish your modified source under the same license. That keeps wigolo open while preventing a closed, hosted fork. See **[SECURITY.md](SECURITY.md)** to report a vulnerability and **[TRADEMARK.md](TRADEMARK.md)** for use of the name. For commercial-licensing questions, reach out.\n\n\u003cdiv align=\"center\"\u003e\n\u003cbr\u003e\n\nwigolo is free and actively maintained, and it's meant to stay that way.\nIf it saves you a metered search bill, a ⭐, a sharp issue, or a **[☕ coffee](https://buymeacoffee.com/knockoutez)** helps keep it sustainable.\n\n\u003csub\u003eBuilt and maintained by \u003ca href=\"https://github.com/KnockOutEZ\"\u003e@KnockOutEZ\u003c/a\u003e · \u003ca href=\"mailto:ktowhid20@gmail.com\"\u003ektowhid20@gmail.com\u003c/a\u003e · \u003ca href=\"https://x.com/yourtowhid\"\u003eX\u003c/a\u003e · \u003ca href=\"https://www.linkedin.com/in/yourtowhid/\"\u003eLinkedIn\u003c/a\u003e\u003c/sub\u003e\n\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FKnockOutEZ%2Fwigolo","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FKnockOutEZ%2Fwigolo","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FKnockOutEZ%2Fwigolo/lists"}