{"id":45139072,"url":"https://github.com/yigitkonur/mcp-researchpowerpack","last_synced_at":"2026-04-04T08:25:17.886Z","repository":{"id":343861804,"uuid":"1175972411","full_name":"yigitkonur/mcp-researchpowerpack","owner":"yigitkonur","description":"The ultimate research MCP toolkit: Reddit mining, web search with CTR aggregation, AI-powered deep research, and intelligent web scraping","archived":false,"fork":false,"pushed_at":"2026-03-12T05:02:12.000Z","size":777,"stargazers_count":1,"open_issues_count":0,"forks_count":1,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-03-12T10:07:42.362Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"TypeScript","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/yigitkonur.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-03-08T12:48:56.000Z","updated_at":"2026-03-12T05:02:15.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/yigitkonur/mcp-researchpowerpack","commit_stats":null,"previous_names":["yigitkonur/mcp-researchpowerpack"],"tags_count":36,"template":false,"template_full_name":null,"purl":"pkg:github/yigitkonur/mcp-researchpowerpack","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yigitkonur%2Fmcp-researchpowerpack","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yigitkonur%2Fmcp-researchpowerpack/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yigitkonur%2Fmcp-researchpowerpack/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yigitkonur%2Fmcp-researchpowerpack/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/yigitkonur","download_url":"https://codeload.github.com/yigitkonur/mcp-researchpowerpack/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yigitkonur%2Fmcp-researchpowerpack/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":31393052,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-04T04:26:24.776Z","status":"ssl_error","status_checked_at":"2026-04-04T04:23:34.147Z","response_time":60,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-agents","deep-research","mcp","mcp-server","model-context-protocol","reddit","research-automation","typescript","web-scraping","web-search"],"created_at":"2026-02-20T00:32:40.025Z","updated_at":"2026-04-04T08:25:17.861Z","avatar_url":"https://github.com/yigitkonur.png","language":"TypeScript","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003ch1 align=\"center\"\u003e🔬 MCP Research Powerpack\u003c/h1\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cstrong\u003eFive research tools for AI assistants — search, scrape, mine Reddit, and synthesize with LLMs.\u003c/strong\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://www.npmjs.com/package/mcp-research-powerpack\"\u003e\u003cimg src=\"https://img.shields.io/npm/v/mcp-research-powerpack.svg?style=flat-square\u0026color=cb3837\" alt=\"npm\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://www.npmjs.com/package/mcp-research-powerpack\"\u003e\u003cimg src=\"https://img.shields.io/npm/dm/mcp-research-powerpack.svg?style=flat-square\u0026color=blue\" alt=\"downloads\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://nodejs.org/\"\u003e\u003cimg src=\"https://img.shields.io/badge/node-%3E%3D20-93450a.svg?style=flat-square\" alt=\"node\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://opensource.org/licenses/MIT\"\u003e\u003cimg src=\"https://img.shields.io/badge/license-MIT-grey.svg?style=flat-square\" alt=\"license\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://modelcontextprotocol.io\"\u003e\u003cimg src=\"https://img.shields.io/badge/MCP-compatible-5a67d8.svg?style=flat-square\" alt=\"MCP\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ccode\u003enpx mcp-research-powerpack\u003c/code\u003e\n\u003c/p\u003e\n\n---\n\nAn [MCP](https://modelcontextprotocol.io) server that gives Claude, Cursor, Windsurf, and any MCP-compatible AI assistant a complete research toolkit. Google search, Reddit deep-dives, web scraping with AI extraction, and multi-model deep research — all as tools that chain into each other.\n\nZero config to start. Each API key you add unlocks more capabilities.\n\n## Tools\n\n| Tool | What it does | Requires |\n|:-----|:-------------|:---------|\n| **`web_search`** | Parallel Google search across 3–100 keywords with CTR-weighted ranking and consensus detection | `SERPER_API_KEY` |\n| **`search_reddit`** | Same search engine filtered to reddit.com — 10–50 queries in parallel | `SERPER_API_KEY` |\n| **`get_reddit_post`** | Fetch 2–50 Reddit posts with full comment trees, smart comment budget allocation | `REDDIT_CLIENT_ID` + `REDDIT_CLIENT_SECRET` |\n| **`scrape_links`** | Scrape 1–50 URLs with JS rendering fallback, HTML→Markdown, optional AI extraction | `SCRAPEDO_API_KEY` |\n| **`deep_research`** | Send questions to research-capable models (Grok, Gemini) with web search, file attachments | `OPENROUTER_API_KEY` |\n\nTools are designed to **chain**: `web_search` → `scrape_links` → `search_reddit` → `get_reddit_post` → `deep_research` for synthesis. Each tool suggests the next logical step in its output.\n\n## Quick Start\n\n### Claude Desktop / Claude Code\n\nAdd to your MCP config (`~/Library/Application Support/Claude/claude_desktop_config.json`):\n\n```json\n{\n  \"mcpServers\": {\n    \"research-powerpack\": {\n      \"command\": \"npx\",\n      \"args\": [\"-y\", \"mcp-research-powerpack\"],\n      \"env\": {\n        \"SERPER_API_KEY\": \"your-key-here\",\n        \"OPENROUTER_API_KEY\": \"your-key-here\"\n      }\n    }\n  }\n}\n```\n\n### Cursor\n\nAdd to `.cursor/mcp.json` in your project:\n\n```json\n{\n  \"mcpServers\": {\n    \"research-powerpack\": {\n      \"command\": \"npx\",\n      \"args\": [\"-y\", \"mcp-research-powerpack\"],\n      \"env\": {\n        \"SERPER_API_KEY\": \"your-key-here\"\n      }\n    }\n  }\n}\n```\n\n### From Source\n\n```bash\ngit clone https://github.com/yigitkonur/mcp-research-powerpack.git\ncd mcp-research-powerpack\npnpm install \u0026\u0026 pnpm build\npnpm start\n```\n\n### HTTP Transport\n\n```bash\nMCP_TRANSPORT=http MCP_PORT=3000 npx mcp-research-powerpack\n```\n\nExposes `/mcp` endpoint (POST/GET/DELETE with session headers) and `/health`.\n\n## API Keys\n\nEach key unlocks a capability. Missing keys silently disable their tools — the server never crashes.\n\n| Variable | Enables | Free Tier |\n|:---------|:--------|:----------|\n| `SERPER_API_KEY` | `web_search`, `search_reddit` | 2,500 searches/mo — [serper.dev](https://serper.dev) |\n| `REDDIT_CLIENT_ID` + `REDDIT_CLIENT_SECRET` | `get_reddit_post` | Unlimited — [reddit.com/prefs/apps](https://www.reddit.com/prefs/apps) (script type) |\n| `SCRAPEDO_API_KEY` | `scrape_links` | 1,000 credits/mo — [scrape.do](https://scrape.do) |\n| `OPENROUTER_API_KEY` | `deep_research`, LLM extraction | Pay-per-token — [openrouter.ai](https://openrouter.ai) |\n| `CEREBRAS_API_KEY` | Cerebras LLM extraction | — |\n| `USE_CEREBRAS` | Enable Cerebras for extraction (set `true`) | `false` |\n\n## Configuration\n\nOptional tuning via environment variables:\n\n| Variable | Default | Description |\n|:---------|:--------|:------------|\n| `RESEARCH_MODEL` | `x-ai/grok-4-fast` | Primary deep research model |\n| `RESEARCH_FALLBACK_MODEL` | `google/gemini-2.5-flash` | Fallback when primary fails |\n| `LLM_EXTRACTION_MODEL` | `openai/gpt-oss-120b:nitro` | Model for scrape/reddit AI extraction |\n| `DEFAULT_REASONING_EFFORT` | `high` | Research depth: `low`, `medium`, `high` |\n| `DEFAULT_MAX_URLS` | `100` | Max search results per research question (10–200) |\n| `API_TIMEOUT_MS` | `1800000` | Request timeout in ms (default: 30 min) |\n| `MCP_TRANSPORT` | `stdio` | Transport mode: `stdio` or `http` |\n| `MCP_PORT` | `3000` | Port for HTTP mode |\n| `USE_CEREBRAS` | `false` | Set to `true` to use Cerebras for extraction instead of OpenRouter |\n| `CEREBRAS_API_KEY` | — | API key for Cerebras cloud — [cloud.cerebras.ai](https://cloud.cerebras.ai) |\n\n### Cerebras Support\n\nWhen `USE_CEREBRAS=true` and `CEREBRAS_API_KEY` are set, the `scrape_links` tool uses Cerebras (Z.ai GLM 4.7) for AI content extraction instead of OpenRouter. This provides:\n\n- **Ultra-fast extraction** — Cerebras inference is optimized for speed\n- **Independent from OpenRouter** — extraction works even without `OPENROUTER_API_KEY`\n- **Automatic fallback** — if Cerebras is not configured, falls back to OpenRouter\n\n```bash\n# Enable Cerebras for extraction\nUSE_CEREBRAS=true CEREBRAS_API_KEY=your-key npx mcp-research-powerpack\n```\n\n### Network Resilience\n\nAll LLM API calls include built-in stability protections:\n\n- **Request deadlines** — hard timeout prevents calls from hanging indefinitely\n- **Stall detection** — if no response arrives within a threshold, the request is aborted and retried\n- **Exponential backoff** — transient failures (429, 5xx) retry with jitter to avoid thundering herd\n- **Connection loss recovery** — network errors (ECONNRESET, ECONNREFUSED) trigger automatic retry\n- **Graceful degradation** — all tools return structured errors instead of crashing\n\n## How It Works\n\n### Search Ranking\n\nResults from multiple queries are deduplicated by normalized URL and scored using **CTR-weighted position values** (position 1 = 100.0, position 10 = 12.56). URLs appearing across multiple queries get a consensus marker. Frequency threshold starts at ≥3, falls back to ≥2, then ≥1 to ensure results.\n\n### Reddit Comment Budget\n\nGlobal budget of **1,000 comments**, max 200 per post. After the first pass, surplus from posts with fewer comments is redistributed to truncated posts in a second fetch pass.\n\n### Scraping Pipeline\n\n**Three-mode fallback** per URL: basic → JS rendering → JS + US geo-targeting. Results go through HTML→Markdown conversion (Turndown), then optional AI extraction with a 100K char input cap and 8,000 token output per URL.\n\n### Deep Research\n\n**32,000 token budget** divided across questions (1 question = 32K, 10 questions = 3.2K each). Gemini models get `google_search` tool access. Grok/Perplexity get `search_parameters` with citations. Primary model fails → automatic fallback to secondary model.\n\n### File Attachments\n\n`deep_research` can read **local files** and include them as context. Files over 600 lines are smart-truncated (first 500 + last 100 lines). Line ranges supported. Line numbers preserved in output.\n\n## Concurrency\n\n| Operation | Parallel Limit |\n|:----------|:---------------|\n| Web search keywords | 8 |\n| Reddit search queries | 8 |\n| Reddit post fetches per batch | 5 (batches of 10) |\n| URL scraping per batch | 10 (batches of 30) |\n| LLM extraction | 3 |\n| Deep research questions | 3 |\n\nAll clients use **manual retry with exponential backoff and jitter**. The OpenAI SDK's built-in retry is disabled (`maxRetries: 0`).\n\n## Architecture\n\n```\nsrc/\n├── index.ts                    Entry point — STDIO + HTTP transport, graceful shutdown\n├── worker.ts                   Cloudflare Workers entry (Durable Objects)\n├── config/\n│   ├── index.ts                Env parsing, capability detection, lazy Proxy config\n│   ├── loader.ts               YAML → Zod → JSON Schema pipeline\n│   └── yaml/tools.yaml         Single source of truth for tool definitions\n├── schemas/                    Zod input validation (deep-research, scrape-links, web-search)\n├── tools/\n│   ├── registry.ts             Tool lookup → capability check → validate → execute\n│   ├── search.ts               web_search handler\n│   ├── reddit.ts               search_reddit + get_reddit_post handlers\n│   ├── scrape.ts               scrape_links handler\n│   └── research.ts             deep_research handler\n├── clients/\n│   ├── search.ts               Google Serper API client\n│   ├── reddit.ts               Reddit OAuth + comment tree parser\n│   ├── scraper.ts              Scrape.do client with fallback modes\n│   └── research.ts             OpenRouter client with model-specific handling\n├── services/\n│   ├── llm-processor.ts        Shared LLM extraction (singleton OpenAI client)\n│   ├── markdown-cleaner.ts     HTML → Markdown via Turndown\n│   └── file-attachment.ts      Local file reading with line ranges\n└── utils/\n    ├── retry.ts                Shared backoff + retry constants\n    ├── concurrency.ts          Bounded parallel execution (pMap, pMapSettled)\n    ├── url-aggregator.ts       CTR-weighted scoring + consensus detection\n    ├── errors.ts               Error classification + structured errors\n    ├── logger.ts               MCP logging protocol\n    └── response.ts             Standardized 70/20/10 output formatting\n```\n\n## Deploy\n\n### Cloudflare Workers\n\n```bash\nnpx wrangler deploy\n```\n\nUses Durable Objects with SQLite storage. YAML-based tool definitions are replaced with inline definitions since there's no filesystem in Workers.\n\n### npm\n\nPublished as [`mcp-research-powerpack`](https://www.npmjs.com/package/mcp-research-powerpack). Binary names: `mcp-research-powerpack`, `research-powerpack-mcp`.\n\n## Development\n\n```bash\npnpm install          # Install dependencies\npnpm dev              # Run with tsx (live TypeScript)\npnpm build            # Compile to dist/\npnpm typecheck        # Type-check without emitting\npnpm start            # Run compiled output\n```\n\n### Testing\n\n```bash\npnpm test:web-search     # Test web search tool\npnpm test:reddit-search  # Test Reddit search\npnpm test:scrape-links   # Test scraping\npnpm test:deep-research  # Test deep research\npnpm test:all            # Run all tests\npnpm test:check          # Check environment setup\n```\n\n## Contributing\n\n1. Fork the repository\n2. Create a feature branch (`git checkout -b feature/amazing-feature`)\n3. Make your changes\n4. Run `pnpm typecheck \u0026\u0026 pnpm build` to verify\n5. Commit (`git commit -m 'feat: add amazing feature'`)\n6. Push to your branch (`git push origin feature/amazing-feature`)\n7. Open a Pull Request\n\n## License\n\n[MIT](https://opensource.org/licenses/MIT) © [Yiğit Konur](https://github.com/yigitkonur)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fyigitkonur%2Fmcp-researchpowerpack","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fyigitkonur%2Fmcp-researchpowerpack","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fyigitkonur%2Fmcp-researchpowerpack/lists"}