{"id":125080,"url":"https://github.com/crawlee-cloud/awesome-web-scraping","name":"awesome-web-scraping","description":"A curated list of web scraping tools, libraries, and resources.","projects_count":43,"last_synced_at":"2026-09-04T18:00:45.222Z","repository":{"id":333123428,"uuid":"1136286220","full_name":"crawlee-cloud/awesome-web-scraping","owner":"crawlee-cloud","description":"A curated list of web scraping tools, libraries, and resources.","archived":false,"fork":false,"pushed_at":"2026-01-17T19:17:52.000Z","size":4,"stargazers_count":2,"open_issues_count":2,"forks_count":2,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-08-16T01:06:20.338Z","etag":null,"topics":["apify","data-mining","scraper","scraping","self-hosted","webscraping"],"latest_commit_sha":null,"homepage":"https://crawlee.cloud","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/crawlee-cloud.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-01-17T12:11:46.000Z","updated_at":"2026-02-17T08:13:38.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/crawlee-cloud/awesome-web-scraping","commit_stats":null,"previous_names":["crawlee-cloud/awesome-web-scraping"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/crawlee-cloud/awesome-web-scraping","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/crawlee-cloud%2Fawesome-web-scraping","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/crawlee-cloud%2Fawesome-web-scraping/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/crawlee-cloud%2Fawesome-web-scraping/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/crawlee-cloud%2Fawesome-web-scraping/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/crawlee-cloud","download_url":"https://codeload.github.com/crawlee-cloud/awesome-web-scraping/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/crawlee-cloud%2Fawesome-web-scraping/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":336071963,"owners_count":37021448,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-04T02:00:06.169Z","response_time":115,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2026-03-17T19:49:23.533Z","updated_at":"2026-09-04T18:00:45.223Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["📝 License","\u003ca name=\"data-extraction\"\u003e\u003c/a\u003e⛏️ Data Extraction","\u003ca name=\"scraping-frameworks\"\u003e\u003c/a\u003e🕷️ Scraping Frameworks","\u003ca name=\"browser-automation\"\u003e\u003c/a\u003e🎭 Browser Automation","\u003ca name=\"scheduling\"\u003e\u003c/a\u003e⏰ Scheduling","\u003ca name=\"anti-detection\"\u003e\u003c/a\u003e🛡️ Anti-Detection","\u003ca name=\"utils--user-agents\"\u003e\u003c/a\u003e🧰 Utils \u0026 User Agents","\u003ca name=\"ai--llm-scraping\"\u003e\u003c/a\u003e🤖 AI \u0026 LLM Scraping","\u003ca name=\"captcha-solving\"\u003e\u003c/a\u003e🧩 CAPTCHA Solving","\u003ca name=\"tutorials\"\u003e\u003c/a\u003e📚 Tutorials","🚀 Self-Hosted Platforms"],"sub_categories":[],"readme":"# Awesome Web Scraping [![Awesome](https://awesome.re/badge.svg)](https://awesome.re)\n\nA curated list of web scraping tools, libraries, and resources.\n\n\u003e Maintained by [Crawlee Cloud](https://crawlee.cloud) — self-hosted web scraping platform.\n\n---\n\n## 📖 Contents\n\n- [AI \u0026 LLM Scraping](#ai--llm-scraping)\n- [Scraping Frameworks](#scraping-frameworks)\n- [Browser Automation](#browser-automation)\n- [Anti-Detection](#anti-detection)\n- [Data Extraction](#data-extraction)\n- [CAPTCHA Solving](#captcha-solving)\n- [Proxy Management](#proxy-management)\n- [Utils \u0026 User Agents](#utils--user-agents)\n- [Scheduling](#scheduling)\n- [Tutorials](#tutorials)\n\n---\n\n## \u003ca name=\"ai--llm-scraping\"\u003e\u003c/a\u003e🤖 AI \u0026 LLM Scraping\n \nTools for extracting data for Large Language Models.\n \n| Tool | Language | Description |\n|------|----------|-------------|\n| [Crawl4AI](https://crawl4ai.com) | Python | Open-source LLM-friendly web crawler and scraper |\n| [Firecrawl](https://firecrawl.dev) | Multi | Turn websites into LLM-ready markdown |\n| [ScrapeGraphAI](https://scrapegraphai.com) | Python | Python library for graph-based AI scraping |\n| [Stagehand](https://stagehand.dev) | TypeScript | AI-powered programmable browser |\n \n---\n \n## \u003ca name=\"scraping-frameworks\"\u003e\u003c/a\u003e🕷️ Scraping Frameworks\n\nFull-featured frameworks for building web scrapers.\n\n| Tool | Language | Description |\n|------|----------|-------------|\n| [Colly](https://go-colly.org) | Go | Fast and elegant scraping framework |\n| [Crawlee](https://crawlee.dev) | TypeScript | Reliable crawling library with autoscaling, session management, and stealth features |\n| [Ferret](https://github.com/MontFerret/ferret) | Go | Declarative web scraping |\n| [Scrapy](https://scrapy.org) | Python | Battle-tested framework for large-scale scraping |\n\n---\n\n## \u003ca name=\"browser-automation\"\u003e\u003c/a\u003e🎭 Browser Automation\n\nHeadless browser control for JavaScript-heavy sites.\n\n| Tool | Language | Description |\n|------|----------|-------------|\n| [Browserbase](https://browserbase.com) | - | Serverless headless browser platform |\n| [Cypress](https://cypress.io) | JavaScript | E2E testing with scraping capabilities |\n| [Playwright](https://playwright.dev) | Multi | Cross-browser automation by Microsoft |\n| [Puppeteer](https://pptr.dev) | JavaScript | Headless Chrome/Chromium control |\n| [rod](https://github.com/go-rod/rod) | Go | High-level Chrome DevTools controller |\n| [Selenium](https://selenium.dev) | Multi | Industry standard browser automation |\n| [Steel](https://steel.dev) | - | Browser API for AI agents |\n\n---\n\n## \u003ca name=\"anti-detection\"\u003e\u003c/a\u003e🛡️ Anti-Detection\n\nTools for avoiding bot detection and CAPTCHAs.\n\n| Tool | Description |\n|------|-------------|\n| [Camoufox](https://camoufox.com) | Stealthy Firefox automation |\n| [curl-impersonate](https://github.com/lwthiker/curl-impersonate) | curl with browser TLS fingerprints |\n| [Rebrowser Patches](https://github.com/rebrowser/rebrowser-patches) | Playwright/Puppeteer anti-detection |\n| [undetected-chromedriver](https://github.com/ultrafunkamsterdam/undetected-chromedriver) | Selenium patch for anti-detection |\n\n---\n\n## \u003ca name=\"data-extraction\"\u003e\u003c/a\u003e⛏️ Data Extraction\n\nHTML parsing and data extraction libraries.\n\n| Tool | Language | Description |\n|------|----------|-------------|\n| [Beautiful Soup](https://www.crummy.com/software/BeautifulSoup/) | Python | HTML/XML parsing |\n| [Cheerio](https://cheerio.js.org) | JavaScript | Fast jQuery-like HTML parsing |\n| [jsdom](https://github.com/jsdom/jsdom) | JavaScript | DOM implementation for Node.js |\n| [lxml](https://lxml.de) | Python | High-performance XML/HTML processing |\n| [Parsel](https://github.com/scrapy/parsel) | Python | XPath/CSS selector extraction |\n| [Selectolax](https://github.com/rushter/selectolax) | Python | Ultra-fast HTML5 parser using Modest engine |\n\n---\n\n## \u003ca name=\"captcha-solving\"\u003e\u003c/a\u003e🧩 CAPTCHA Solving\n\nServices and libraries to solve CAPTCHAs.\n\n| Tool | Type | Description |\n|------|------|-------------|\n| [2Captcha](https://2captcha.com) | Service | Human-powered CAPTCHA solving service |\n| [Anti-Captcha](https://anti-captcha.com) | Service | Reliable CAPTCHA solving API |\n| [CapMonster Cloud](https://capmonster.cloud) | Service | AI-powered cloud CAPTCHA solving service |\n| [CapSolver](https://capsolver.com) | Service | AI-powered CAPTCHA solving |\n| [nocaptchaai](https://nocaptchaai.com) | Service | AI solution for recaptcha/hcaptcha |\n\n---\n\n## \u003ca name=\"proxy-management\"\u003e\u003c/a\u003e🌐 Proxy Management\n\nRotating proxies and IP management.\n\n| Type | Description |\n|------|-------------|\n| **Residential** | Real user IPs, higher trust |\n| **Datacenter** | Fast, cheap, easily detected |\n| **Mobile** | 4G/5G IPs, highest trust |\n| **ISP** | Static residential IPs |\n\n\u003e Popular providers: Bright Data, Oxylabs, Smartproxy, IPRoyal\n\n---\n\n## \u003ca name=\"utils--user-agents\"\u003e\u003c/a\u003e🧰 Utils \u0026 User Agents\n\nHelper libraries for common scraping tasks.\n\n| Tool | Language | Description |\n|------|----------|-------------|\n| [fake-useragent](https://github.com/fake-useragent/fake-useragent) | Python | Random User-Agent generator |\n| [protego](https://github.com/scrapy/protego) | Python | Pure-Python robots.txt parser |\n| [robots-parser](https://github.com/samclarke/robots-parser) | JavaScript | robots.txt parser for Node.js |\n| [user-agents](https://github.com/intoli/user-agents) | JavaScript | Comprehensive User-Agent generator |\n\n---\n\n## \u003ca name=\"scheduling\"\u003e\u003c/a\u003e⏰ Scheduling\n\nJob scheduling for recurring scrapes.\n\n| Tool | Language | Description |\n|------|----------|-------------|\n| [APScheduler](https://apscheduler.readthedocs.io) | Python | Advanced scheduler |\n| [BullMQ](https://docs.bullmq.io) | JavaScript | Redis-based job queue |\n| [Celery](https://docs.celeryq.dev) | Python | Distributed task queue |\n| [node-cron](https://github.com/node-cron/node-cron) | JavaScript | Cron-like scheduler |\n\n---\n\n## \u003ca name=\"tutorials\"\u003e\u003c/a\u003e📚 Tutorials\n\nLearning resources for web scraping.\n\n- [Apify Academy](https://docs.apify.com/academy) — Free web scraping course\n- [Crawlee Docs](https://crawlee.dev/docs) — Official Crawlee documentation\n- [ScrapingBee Blog](https://www.scrapingbee.com/blog/) — Practical guides\n\n---\n\n## 🚀 Self-Hosted Platforms\n\nRun scrapers on your own infrastructure.\n\n| Platform | Description |\n|----------|-------------|\n| [Crawlee Cloud](https://crawlee.cloud) | Open-source, self-hosted Actor platform |\n\n---\n\n## 🤝 Contributing\n\nContributions welcome! Please read the [contributing guidelines](CONTRIBUTING.md) first.\n\n---\n\n## 📝 License\n\n[![CC0](https://licensebuttons.net/p/zero/1.0/88x31.png)](https://creativecommons.org/publicdomain/zero/1.0/)\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/crawlee-cloud%2Fawesome-web-scraping/projects"}