{"id":51718064,"url":"https://github.com/PKHarsimran/website-downloader","last_synced_at":"2026-07-28T16:00:29.571Z","repository":{"id":251340452,"uuid":"823845627","full_name":"PKHarsimran/website-downloader","owner":"PKHarsimran","description":"Website-downloader is a powerful and versatile Python script designed to download entire websites along with all their assets. This tool allows you to create a local copy of a website, including HTML pages, images, CSS, JavaScript files, and other resources. It is ideal for web archiving, offline browsing, and web development.","archived":false,"fork":false,"pushed_at":"2026-07-06T21:32:18.000Z","size":230,"stargazers_count":140,"open_issues_count":2,"forks_count":37,"subscribers_count":4,"default_branch":"main","last_synced_at":"2026-07-06T22:17:03.862Z","etag":null,"topics":["automation","beautifulsoup","data-mining","html","internet-tools","offline-browsing","open-source","python","python-scripts","requests","web-archiving","web-scraping","website-cloner","website-downloader","wget"],"latest_commit_sha":null,"homepage":"https://harsim.ca/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/PKHarsimran.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2024-07-03T20:59:45.000Z","updated_at":"2026-07-06T21:32:20.000Z","dependencies_parsed_at":"2026-06-28T17:02:51.252Z","dependency_job_id":null,"html_url":"https://github.com/PKHarsimran/website-downloader","commit_stats":null,"previous_names":["pkharsimran/website-downloader"],"tags_count":11,"template":false,"template_full_name":null,"purl":"pkg:github/PKHarsimran/website-downloader","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKHarsimran%2Fwebsite-downloader","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKHarsimran%2Fwebsite-downloader/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKHarsimran%2Fwebsite-downloader/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKHarsimran%2Fwebsite-downloader/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/PKHarsimran","download_url":"https://codeload.github.com/PKHarsimran/website-downloader/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PKHarsimran%2Fwebsite-downloader/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35999100,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-28T02:00:06.341Z","response_time":109,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["automation","beautifulsoup","data-mining","html","internet-tools","offline-browsing","open-source","python","python-scripts","requests","web-archiving","web-scraping","website-cloner","website-downloader","wget"],"created_at":"2026-07-17T06:00:35.699Z","updated_at":"2026-07-28T16:00:29.555Z","avatar_url":"https://github.com/PKHarsimran.png","language":"Python","funding_links":["https://www.paypal.com/donate/?business=PJVPSXG6V4CUG\u0026no_recurring=1\u0026item_name=Thank+you+for+the+coffee+%3A%29\u0026currency_code=CAD"],"categories":["Projects"],"sub_categories":["Web Scraping"],"readme":"\u003cdiv align=\"center\"\u003e\n\n# 🌐 Website Downloader CLI\n\n**Turn any website you're authorized to copy into a fast, browsable offline mirror — with one command.**\n\n[![CI - Website Downloader](https://github.com/PKHarsimran/website-downloader/actions/workflows/python-app.yml/badge.svg)](https://github.com/PKHarsimran/website-downloader/actions/workflows/python-app.yml)\n[![Lint \u0026 Style](https://github.com/PKHarsimran/website-downloader/actions/workflows/lint.yml/badge.svg)](https://github.com/PKHarsimran/website-downloader/actions/workflows/lint.yml)\n[![Python](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n[![Last Commit](https://img.shields.io/github/last-commit/PKHarsimran/website-downloader)](https://github.com/PKHarsimran/website-downloader/commits/main)\n[![PRs Welcome](https://img.shields.io/badge/PRs-welcome-brightgreen.svg)](https://github.com/PKHarsimran/website-downloader/pulls)\n\n*A modern, hackable alternative to `wget --mirror` and HTTrack — built in pure Python, without dragging in a heavy crawler framework.*\n\n📖 **[Read the Wiki](https://github.com/PKHarsimran/website-downloader/wiki)** — full CLI reference, cookbook, and troubleshooting · 📦 **[Install from PyPI](https://pypi.org/project/website-downloader/)**\n\n\u003c/div\u003e\n\n---\n\n\u003cdiv align=\"center\"\u003e\n\n![Website Downloader CLI in action](docs/demo.gif)\n\n\u003c/div\u003e\n\nOpen `example_backup/index.html` in your browser — the whole site works from disk: pages, styles, scripts, images, fonts, and media, all with links rewritten for offline browsing.\n\n## ⚡ Quick Start\n\n```bash\ngit clone https://github.com/PKHarsimran/website-downloader.git\ncd website-downloader\n\npython -m venv .venv\n.venv\\Scripts\\activate        # macOS/Linux: source .venv/bin/activate\npip install -e .\n\nwebsite-downloader --url https://example.com --destination example_backup --max-pages 100\n```\n\nThe classic script entry point still works too:\n\n```bash\npython website-downloader.py --url https://example.com --destination example_backup\n```\n\n## 🤔 Why Not Just wget or HTTrack?\n\nThose tools are great — until you hit a modern website. This project exists for the gap between \"one-liner that misses half the assets\" and \"write your own Scrapy project.\"\n\n| | **website-downloader** | wget --mirror | HTTrack | Scrapy |\n| --- | :-: | :-: | :-: | :-: |\n| Modern assets: `srcset`, `data-src`, `poster`, CSS `@import`, JS asset strings | ✅ | partial | partial | build it yourself |\n| JavaScript rendering (React, Vue, Next.js) | ✅ Playwright | ❌ | ❌ | plugin |\n| Incremental re-mirroring (`ETag` / `Last-Modified`) | ✅ | timestamps only | ✅ | manual |\n| Cookies + custom headers for authorized portals | ✅ | ✅ | ✅ | ✅ |\n| Selective CDN mirroring with a domain allowlist | ✅ | ❌ | partial | manual |\n| Zip + WARC export | ✅ | WARC ✅ | ❌ | manual |\n| Windows-safe paths (long paths, reserved names, query hashing) | ✅ | ❌ | partial | manual |\n| Small, readable Python codebase you can extend | ✅ | ❌ (C) | ❌ (C) | framework |\n\n## ✨ Highlights\n\n- 🚀 **Fast** — parallel page fetching (`--page-threads`, ~3–4× faster on multi-page sites), threaded asset downloads, and optional lxml parsing (`pip install -e \".[fast]\"`).\n- 🔁 **Incremental** — `--update` skips unchanged pages and assets using `ETag`/`Last-Modified`, perfect for recurring archives.\n- ⚛️ **JavaScript-aware** — optional Playwright rendering for client-rendered sites (`--render-js`).\n- 🍪 **Authenticated** — reuse browser cookies and custom headers for portals, intranets, and staging sites you're allowed to access.\n- 🧭 **Sitemap seeding** — start from `sitemap.xml` (including nested sitemap indexes) for complete discovery.\n- 📦 **Portable output** — export mirrors as zip archives or WARC 1.1 response records.\n- 🤝 **Polite by default** — sequential pages unless you opt in, `--respect-robots`, `--delay`, retry with backoff, and per-asset size caps.\n- 🪟 **Cross-platform paths** — sanitizes Windows reserved names, shortens long paths, and hashes query strings to avoid collisions.\n- 🧪 **Tested** — pytest suite running against a real local HTTP fixture server, with CI and lint gates.\n\n## 🛠 How It Works\n\n```mermaid\nflowchart TD\n    A[\"Start with a URL and CLI options\"] --\u003e B[\"Create session with cookies, headers, retries\"]\n    B --\u003e C{\"Use sitemap?\"}\n    C -- \"Yes\" --\u003e D[\"Load sitemap URLs into the page queue\"]\n    C -- \"No\" --\u003e E[\"Queue the starting URL\"]\n    D --\u003e F[\"Fetch next page\"]\n    E --\u003e F\n    F --\u003e G{\"Update cache says unchanged?\"}\n    G -- \"Yes\" --\u003e H[\"Reuse saved local file\"]\n    G -- \"No\" --\u003e I{\"Render JavaScript?\"}\n    I -- \"No\" --\u003e J[\"Download HTML with requests\"]\n    I -- \"Yes\" --\u003e K[\"Render page with Playwright\"]\n    J --\u003e L[\"Parse HTML with BeautifulSoup\"]\n    K --\u003e L\n    H --\u003e L\n    L --\u003e M[\"Find page links and asset links\"]\n    M --\u003e N{\"Same-site page?\"}\n    N -- \"Yes\" --\u003e O[\"Queue page for crawling\"]\n    N -- \"No\" --\u003e P{\"Asset allowed?\"}\n    P -- \"Yes\" --\u003e Q[\"Download asset\"]\n    P -- \"No\" --\u003e R[\"Keep original reference or skip\"]\n    O --\u003e S[\"Rewrite links for offline browsing\"]\n    Q --\u003e S\n    R --\u003e S\n    S --\u003e T[\"Save mirror folder\"]\n    T --\u003e U{\"Export requested?\"}\n    U -- \"Zip/WARC\" --\u003e V[\"Write portable archive\"]\n    U -- \"No\" --\u003e W[\"Open index.html locally\"]\n    V --\u003e W\n```\n\nIn plain English:\n\n1. You give the CLI a starting URL and optional crawl settings.\n2. It can seed pages from `sitemap.xml`, custom headers, cookies, and robots rules.\n3. It downloads or optionally renders each page with Playwright.\n4. It finds links, images, scripts, stylesheets, fonts, media, and metadata assets.\n5. It follows same-site pages up to your `--max-pages` limit.\n6. It saves assets locally and rewrites references so pages still work offline.\n7. With `--update`, unchanged resources are skipped using cache metadata.\n8. With `--zip-output` or `--warc-output`, the result is also exported as an archive.\n\n## 📦 Install Options\n\nStart with the core install, then add extras only when you need them:\n\n| Install | Use when you want |\n| --- | --- |\n| `pip install -e .` | Normal static-site crawling with `requests` and BeautifulSoup. |\n| `pip install -e \".[fast]\"` | Faster HTML parsing with lxml (used automatically when installed). |\n| `pip install -e \".[render]\"` | Playwright-powered JavaScript rendering with `--render-js` or `--headless`. |\n| `pip install -e \".[ux]\"` | Rich-powered terminal progress with `--progress`. |\n| `pip install -e \".[dev]\"` | Tests, formatting, linting, and local contributor work. |\n\n## 📖 Cookbook\n\nMirror a small public site:\n\n```bash\nwebsite-downloader ^\n  --url https://example.com ^\n  --destination example_backup ^\n  --max-pages 50\n```\n\nSpeed up a large mirror with parallel page fetching:\n\n```bash\nwebsite-downloader ^\n  --url https://example.com ^\n  --max-pages 500 ^\n  --page-threads 4\n```\n\n`--page-threads` defaults to 1 so crawls stay polite; raise it only for sites that can handle concurrent requests. `--render-js` always uses a single page worker.\n\nDownload selected CDN assets:\n\n```bash\nwebsite-downloader ^\n  --url https://example.com ^\n  --destination example_backup ^\n  --download-external-assets ^\n  --external-domains cdn.example.com fonts.gstatic.com\n```\n\nMirror an authorized site with cookies:\n\n```bash\nwebsite-downloader ^\n  --url https://intranet.example.com ^\n  --destination intranet_backup ^\n  --cookie-file example-cookie.txt\n```\n\nCookie files use normal cookie header syntax:\n\n```text\nsessionid=abc123; csrftoken=xyz789\n```\n\nSend custom headers such as bearer tokens:\n\n```bash\nwebsite-downloader ^\n  --url https://docs.example.com ^\n  --destination docs_backup ^\n  --header \"Authorization: Bearer \u003ctoken\u003e\" ^\n  --header \"X-Environment: staging\"\n```\n\nUse a sitemap as the crawl seed:\n\n```bash\nwebsite-downloader ^\n  --url https://example.com ^\n  --destination example_backup ^\n  --sitemap\n```\n\nPoint at a custom sitemap URL or local sitemap file:\n\n```bash\nwebsite-downloader --url https://example.com --sitemap https://example.com/sitemap.xml\n```\n\nUse safer crawl limits:\n\n```bash\nwebsite-downloader ^\n  --url https://example.com ^\n  --max-pages 50 ^\n  --threads 4 ^\n  --delay 0.25 ^\n  --respect-robots ^\n  --max-asset-bytes 25000000 ^\n  --user-agent \"WebsiteDownloader/0.2\"\n```\n\nUpdate an existing mirror without re-downloading unchanged resources:\n\n```bash\nwebsite-downloader ^\n  --url https://example.com ^\n  --destination example_backup ^\n  --update\n```\n\nExport a portable zip and WARC archive:\n\n```bash\nwebsite-downloader ^\n  --url https://example.com ^\n  --destination example_backup ^\n  --zip-output example_backup.zip ^\n  --warc-output example_backup.warc\n```\n\n## ⚛️ JavaScript-Rendered Sites\n\nSome modern sites do not expose their real links and assets until JavaScript runs. For those, install the optional Playwright extra:\n\n```bash\npip install -e \".[render]\"\nplaywright install chromium\nwebsite-downloader --url https://example.com --render-js --max-pages 20\n```\n\n`--headless` is also available as a friendly alias for `--render-js`.\n\n`--render-js` and `--headless` are optional because Playwright is heavier than the default `requests` + BeautifulSoup path. Use them when a normal crawl only captures an empty app shell or misses important client-rendered links.\n\n## 📊 Live Progress\n\nInstall the optional UX extra for a Rich-powered terminal dashboard:\n\n```bash\npip install -e \".[ux]\"\nwebsite-downloader --url https://example.com --progress\n```\n\nIf `rich` is not installed, the crawler falls back to normal logging instead of failing.\n\n## 🎛 Feature Flags At A Glance\n\n| Flag | What it does | Best for |\n| --- | --- | --- |\n| `--page-threads` | Fetches HTML pages concurrently (default 1). | Faster mirroring of large sites that tolerate concurrent requests. |\n| `--render-js` / `--headless` | Uses Playwright before parsing the page. | React, Vue, Angular, Next.js, and other client-rendered sites. |\n| `--cookie-file` | Sends saved browser/session cookies. | Authorized portals, staging sites, docs behind login. |\n| `--header` | Adds custom request headers. | Bearer tokens, staging headers, API gateway headers. |\n| `--update` | Reuses cache metadata and skips unchanged resources when the server supports it. | Recurring mirrors and archives. |\n| `--sitemap` | Seeds the crawl from `sitemap.xml` or a supplied sitemap. | Faster, more complete discovery. |\n| `--progress` | Shows a Rich terminal progress dashboard when installed. | Long crawls where visibility matters. |\n| `--zip-output` | Exports the mirror folder as a zip. | Sharing, attaching, or storing snapshots. |\n| `--warc-output` | Writes a simple WARC response archive. | Archival workflows and future replay tooling. |\n\n## 🔁 What Gets Rewritten\n\n| Source | Rewritten for offline use |\n| --- | --- |\n| Page links | `\u003ca href\u003e` for same-site pages |\n| Images and media | `src`, `data-src`, `poster`, `srcset` |\n| Stylesheets and icons | `\u003clink href\u003e` for fetchable resource types |\n| Metadata images | `og:image`, `twitter:image` |\n| Inline styles | `style=\"background: url(...)\"` |\n| CSS files | `url(...)` and `@import` |\n| JavaScript files | Common static asset strings like `/img/logo.png` |\n| External assets | Optional CDN copies under `cdn/\u003cdomain\u003e/...` |\n\nWhen external scripts or stylesheets are localized, the tool removes `integrity` and `crossorigin` where needed because those attributes often break offline copies.\n\n## 📁 Output Example\n\n```text\nexample_backup/\n  index.html\n  about.html\n  assets/\n    site.css\n    app.js\n  img/\n    logo.png\n    hero.webp\n  fonts/\n    inter.woff2\n  cdn/\n    cdn.example.com/\n      library.js\n```\n\nOpen `index.html` in your browser to browse the mirrored copy.\n\n## 🧑‍💻 Development\n\n```bash\npip install -e \".[dev]\"\npytest\nblack . --check\nisort . --check-only\nruff check .\n```\n\nUsing PyCharm? Open the repo folder, point it at a Python 3.10+ virtualenv, run `pip install -e \".[dev]\"` in the terminal, and use the pytest runner on the `tests` folder.\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eProject structure\u003c/b\u003e\u003c/summary\u003e\n\n| Path | Purpose |\n| --- | --- |\n| `website_downloader/cli.py` | Argument parsing, validation, logging, and CLI entry point. |\n| `website_downloader/crawler.py` | Crawl coordination, page/asset worker pools, robots.txt support, and stats. |\n| `website_downloader/http.py` | Requests sessions, HTML fetches, binary downloads, and downloaded CSS/JS post-processing. |\n| `website_downloader/rewrite.py` | HTML, CSS, JavaScript, and `srcset` reference rewriting. |\n| `website_downloader/paths.py` | Filesystem-safe page, asset, and CDN path mapping. |\n| `website_downloader/render.py` | Optional Playwright page rendering. |\n| `website_downloader/cache.py` | Update-mode metadata for `ETag` and `Last-Modified`. |\n| `website_downloader/sitemap.py` | Sitemap and sitemap-index loading. |\n| `website_downloader/progress.py` | Optional Rich progress dashboard. |\n| `website_downloader/exports.py` | Zip and WARC export helpers. |\n| `tests/` | Local pytest suite with a tiny fixture HTTP server. |\n\n\u003c/details\u003e\n\n## 🗺 Roadmap\n\n- `--manifest crawl.json` with pages, assets, status codes, titles, headings, and errors.\n- Login-flow recording for complex SSO sites.\n- Stronger WARC metadata and replay compatibility.\n- Visual diff mode for migration and redesign checks.\n\nHave an idea? [Open an issue](https://github.com/PKHarsimran/website-downloader/issues) — feature requests and bug reports are very welcome.\n\n## 🛡 Responsible Use\n\nOnly mirror sites you own, have permission to archive, or are legally allowed to access. Authentication cookies can expose private content, so keep cookie files out of source control and avoid sharing generated mirrors that contain private data. Use `--respect-robots`, lower `--threads`, and `--delay` for polite crawling.\n\n## 🤝 Contributing\n\nContributions are welcome! Open an issue or pull request for bug reports, feature ideas, or improvements. The codebase is intentionally small and modular — most features live in a single focused module, so it's an easy project to hack on.\n\nBy submitting a contribution, you agree to the [Contributor License Agreement](CLA.md): you keep the copyright to your work and license it to the project so it can be maintained and distributed (including under future licensing terms).\n\nIf this tool saved you time, consider **starring the repo** ⭐ — it helps others find it.\n\n## ☕ Support This Project\n\n[Donate via PayPal](https://www.paypal.com/donate/?business=PJVPSXG6V4CUG\u0026no_recurring=1\u0026item_name=Thank+you+for+the+coffee+%3A%29\u0026currency_code=CAD)\n\n## 📄 License \u0026 Ownership\n\nThis project is released under the [MIT License](LICENSE) — free to use, fork, modify, and ship, including in commercial products, as long as the copyright and license notice are kept.\n\nThe copyright is owned by Harsimran Sidhu. The MIT license grants broad permission to use the code; it does not transfer ownership. The name \"Website Downloader CLI\" and any associated branding are not covered by the code license.\n\n### Commercial use, hosting, and support\n\nThe open-source license already covers most commercial use. If you'd like something the MIT license doesn't provide — a **commercial/OEM license** with different terms, a **hosted or managed version**, **priority support**, or **custom features** — reach out via a [GitHub issue](https://github.com/PKHarsimran/website-downloader/issues) or the contact on the maintainer's [GitHub profile](https://github.com/PKHarsimran).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FPKHarsimran%2Fwebsite-downloader","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FPKHarsimran%2Fwebsite-downloader","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FPKHarsimran%2Fwebsite-downloader/lists"}