{"id":29652,"url":"https://github.com/all-the-data/awesome-data-hoarding","name":"awesome-data-hoarding","description":"How to save everything online. Tools for for scraping, saving, downloading, hoarding, archiving, etc.","projects_count":44,"last_synced_at":"2026-07-28T17:00:23.706Z","repository":{"id":73756501,"uuid":"460793914","full_name":"all-the-data/awesome-data-hoarding","owner":"all-the-data","description":"How to save everything online. Tools for for scraping, saving, downloading, hoarding, archiving, etc.","archived":false,"fork":false,"pushed_at":"2026-05-29T11:46:41.000Z","size":123,"stargazers_count":45,"open_issues_count":0,"forks_count":6,"subscribers_count":1,"default_branch":"main","last_synced_at":"2026-07-09T19:04:04.629Z","etag":null,"topics":["archiving","awesome","awesome-list","data-hoarder","hoarding","reddit","reddit-downloader"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"cc-by-sa-4.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/all-the-data.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2022-02-18T09:44:20.000Z","updated_at":"2026-06-25T02:48:23.000Z","dependencies_parsed_at":"2024-03-06T05:59:11.937Z","dependency_job_id":"ff20af0a-5b8f-409d-8b83-2773fbb427a6","html_url":"https://github.com/all-the-data/awesome-data-hoarding","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/all-the-data/awesome-data-hoarding","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/all-the-data%2Fawesome-data-hoarding","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/all-the-data%2Fawesome-data-hoarding/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/all-the-data%2Fawesome-data-hoarding/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/all-the-data%2Fawesome-data-hoarding/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/all-the-data","download_url":"https://codeload.github.com/all-the-data/awesome-data-hoarding/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/all-the-data%2Fawesome-data-hoarding/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36001047,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-28T02:00:06.341Z","response_time":109,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-01-13T12:58:00.681Z","updated_at":"2026-07-28T17:00:23.706Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Processing tools","Quick reference","Scraping tools","Techniques"],"sub_categories":["General purpose","Extract playlist data from YouTube and YT Music","Case studies"],"readme":"# awesome-data-hoarding\n\nA concise cheat-sheet of commands and tools for scraping, saving, hoarding, archiving, collecting, organising and browsing data.\n\nInspired by Reddit's [/r/DataHoarder](https://www.reddit.com/r/DataHoarder/)\n\n## Quick reference\n\nWhich archiving tool should you choose for each web service?\n\n- Amazon orders: [amazon-orders](https://github.com/alexdlaird/amazon-orders)\n\n- Amazon Video: Unknown. Check torrents instead.\n\n- BBC iPlayer: [youtube-dl](https://youtube-dl.org/) / [yt-dlp](https://github.com/yt-dlp/yt-dlp)\n\n- Discord: DiscordChatExporter (see below for notes)\n\n- Mediawiki website: Native dump using `/wiki/Special:AllPages` and `/wiki/Special:Export`.\n\n- Netflix: Unknown. Check torrents instead.\n\n- Reddit: Various tools\n  - Tools to save whole threads\n    - [Bulk-Downloader-For-Reddit](https://github.com/aliparlakci/bulk-downloader-for-reddit)\n    - [BDFRX](https://github.com/OMEGARAZER/bulk-downloader-for-reddit-x#differences-from-bdfr)\n    - [Gallery-DL](https://github.com/mikf/gallery-dl)\n    - [RipMe](https://github.com/RipMeApp2/ripme).\n  - \"Print\" method for threads\n    - Change `www.reddit.com` to `old.reddit.com` -- all comments will now be expanded\n    - Sort by: New\n    - Use [cleanly print](https://chromewebstore.google.com/detail/cleanly-print/afloocnncgjhdlacbejppjepboilajdg) chrome extension\n    - Click to remove areas, also click to 'tag' areas for printing.\n  - Historial data dumps: [the-eye](https://the-eye.eu/redarcs/) / [torrents](https://academictorrents.com/userdetails.php?id=9863)\n\n- SoundCloud: [youtube-dl](https://youtube-dl.org/) / [yt-dlp](https://github.com/yt-dlp/yt-dlp) / [lucida.to]([url](https://lucida.to/))\n\n- Tumblr: [TumblThreeApp](https://github.com/TumblThreeApp/TumblThree) (Windows). Viewers: [1](https://github.com/jacob-pro/tumbl-three-viewer), [2](https://github.com/willsheppard/random-scripts/blob/master/TumblThree_BackupViewer.html).\n\n- Twitter: [ThreadReaderApp](https://threadreaderapp.com/)\n\n- Torrents: Use [unblockit](https://www.google.com/search?q=unblockit) for a list of torrent sites. Official [Twitter](https://twitter.com/thepirateproxy) / [Reddit](https://www.reddit.com/r/Unblockit/).\n\n- Private torrent trackers: Might contain any TV or movie ever broadcat. It can be difficult to get an invite, and you may need to maintain an upload ratio.\n\n- Individual web pages:\n  - Save as | Web Page, HTML Only\n  - Save as | Web Page, Single File\n  - Save as | Web Page, Complete\n  - Print | Save as PDF\n  - Chrome extension [SingleFile](https://chromewebstore.google.com/detail/singlefile/mpiodijhokgodhhofbcjdecpffjipkle) \u003c-- Recommended!\n- Websites generally: wget, httrack, [ArchiveBot](https://wiki.archiveteam.org/index.php?title=ArchiveBot) or [Wayback-Archive](https://github.com/GeiserX/Wayback-Archive).\n\n- Images:\n  - There are many good chrome extensions, for example [download-all-images](https://chromewebstore.google.com/detail/download-all-images/nnffbdeachhbpfapjklmpnmjcgamcdmm)\n  - [Greenshot](https://getgreenshot.org/) (Windows)\n  - Screenshot app (Shift-Command-5) (Mac)\n\n- Youtube video/music: [youtube-dl](https://youtube-dl.org/) (see below for notes) / [yt-dlp](https://github.com/yt-dlp/yt-dlp)\n\n- Radio scrobbling / Music identification: [Shazam](https://chromewebstore.google.com/detail/shazam-find-song-names-fr/mmioliijnhnoblpgimnlajmefafdfilb) or [AHA Music finder](https://chromewebstore.google.com/detail/aha-music-song-finder-for/dpacanjfikmhoddligfbehkpomnbgblf)\n\n\n## Scraping tools\n\n- Radio scrobbling\n  - Play radio station with low quality playlist: [La Mega, Malaga](https://onlineradiobox.com/es/lamegaradio/).\n  - Install chrmoe browser extension [Shazam](https://chromewebstore.google.com/detail/shazam-find-song-names-fr/mmioliijnhnoblpgimnlajmefafdfilb) or [AHA Music finder](https://chromewebstore.google.com/detail/aha-music-song-finder-for/dpacanjfikmhoddligfbehkpomnbgblf)\n  - On Linux use `xdotool` to automate clicking on chrome browser extension icons to activate music identification: `watch \"xdotool mousemove 3442 90 click 1; sleep 20; xdotool mousemove 3476 90 click 1; sleep 20\"` (adjust coords as needed)\n  - Does not require speakers to be on\n\n\nDetails of precise sets of commands.\n\n- [wget](https://www.gnu.org/software/wget/manual/wget.html) for websites\n\n```\nwget \\\n    -e 'robots=off' \\\n    --accept '*.*' \\\n    --mirror \\\n    --wait 2 \\\n    --random-wait \\\n    --convert-links \\\n    --user-agent 'Mozilla/5.0 (Windows NT 10.0) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.7113.93 Safari/537.36' \\\n    'http://www.example.com/'\n```\n\n- [Wayback-Archive](https://github.com/GeiserX/Wayback-Archive) for downloading complete websites from the Wayback Machine with full asset preservation for offline viewing (Python, GPL-3.0)\n\n- [StreamRipper](http://streamripper.sourceforge.net) for music\n    - Example: `streamripper ###URL### -u \"FreeAmp/2.x\" -q -l 86400`\n\n- [Chrome DevTools](https://developer.chrome.com/docs/devtools) for anything via a web browser\n    - [network tab](https://developer.chrome.com/docs/devtools/network/reference)\n    - [resources tab](https://developer.chrome.com/docs/devtools/resources)\n\n- Mediawiki for wiki sites\n    - For an XML dump containing wikitext...\n    - Copy names of pages from `/wiki/Special:AllPages`...\n    - Paste into `/wiki/Special:Export`\n    - (optional) Parse resulting wikitext with [mwparserfromhell](https://github.com/earwig/mwparserfromhell).\n\n- [youtube-dl](https://yt-dl.org) / [yt-dlp](https://github.com/yt-dlp/yt-dlp) for Youtube and other video/audio\n  - Video\n\n```\nyt-dlp \\\n    --ignore-errors \\\n    --format 'bestvideo[ext=mp4]+bestaudio[ext=m4a]/best[ext=mp4]/best' \\\n    --output \"%(playlist_title)s/%(title)s.%(ext)s\" \\\n    --throttled-rate 10K \\\n    ###URL###\n```\n\n  - Audio\n\n```\nyt-dlp \\\n    --ignore-errors \\\n    --extract-audio \\\n    --audio-quality 0 \\\n    --audio-format mp3 \\\n    --prefer-ffmpeg \\\n    --output \"%(playlist_title)s/%(artist)s - %(title)s.%(ext)s\" \\\n    --throttled-rate 10K \\\n    ###URL###\n```\n\n  - Audio album playlist\n\n```\nyt-dlp \\\n    ...etc... \\\n    --output \"%(artist)s - %(album)s/%(artist)s - %(album)s - %(playlist_index)02d - %(track)s.%(ext)s\" \\\n    ###URL###\n```\n\n  - Video playlist\n\n```\nyt-dlp \\\n    ...etc... \\\n    --output \"%(playlist_title)s/%(playlist_index)03d - %(artist)s - %(title)s.%(ext)s\" \\\n    ###URL###\n```\n\n  - Multiple playlists\n\n```\nfor URL in $(cat list)\ndo\n    yt-dlp ...etc... \"$URL\"\ndone\n```\n\n- [DiscordChatExporter](https://github.com/Tyrrrz/DiscordChatExporter) + excellent [wiki](https://github.com/Tyrrrz/DiscordChatExporter/wiki)\n    - Example: `docker run --rm -v /var/www/zaphod/adhd:/app/out tyrrrz/discordchatexporter:stable export --channel ###ID### --token ###SECRET### --format Json`\n    - List guilds: `docker run tyrrrz/discordchatexporter:stable guilds`\n    - List channels: `docker run tyrrrz/discordchatexporter:stable channels --guild ###ID###`\n\n## Processing tools\n\n- [jq](https://stedolan.github.io/jq/)\n    - Example: `jq -j -M --stream -f discord1.jq` [discord1.jq](https://gist.github.com/willsheppard/f9b7cc9b130784ffd7bd8f144cf892f8)\n\n- [XPath Helper](https://chrome.google.com/webstore/detail/xpath-helper/hgimnogjllphhhkhlmebbmlgjoejdpjl)\n    - Example: Ctrl-Shift-X (or Command-Shift-X on Mac)\n\n- HAR recorders (if for some reason Chrome's \"Save as HAR\" feature isn't sufficient)\n  - [AutoHAR](https://github.com/Aloisius/autohar)\n  - [HAR Recorder](https://chrome.google.com/webstore/detail/har-recorder/emfabjnfjiknifjlfpjobbecfepplhkd)\n\n- HAR extractors (to retrieve the original content from inside a HAR file) \n  - [How can I extract the contents of a .HAR file](https://www.reddit.com/r/browsers/comments/bczaiz/how_can_i_extract_the_contents_of_a_har_file_in)\n  - [quentint/har-extract.js](https://gist.github.com/quentint/7236a6a7cae187507b0c1dfd4b1ed1c5)\n  - [JC3/harextract](https://github.com/JC3/harextract) + see [forks](https://github.com/JC3/harextract/forks)\n  - [crazypatoto/MorkVideoExtractor](https://github.com/crazypatoto/MorkVideoExtractor)\n  - [outersky/har-tools](https://github.com/outersky/har-tools)\n\n- Convert between file formats\n  - Convert media online: [https://cloudconvert.com/ogg-to-mp3]([url](https://cloudconvert.com/ogg-to-mp3))\n  - Convert media offline: [ffmpeg](https://www.ffmpeg.org/)\n  - Convert documents: [imagemagick]([url](https://imagemagick.org/)) / Pandoc\n\n## Techniques\n\n### Combine streamed .ts files and m3u8 playlist/chunklist into an mpeg/mp4 video\n\n- After extracting the .m4u8 and .ts files from HAR, run something like:\n  - `ffmpeg -i playlist.m3u8 -c copy -bsf:a aac_adtstoasc output.mp4`\n\n### Extract playlist data from YouTube and YT Music\n\nInput: https://music.youtube.com/library/playlists\nGoal: Extract a list of playlists suitable for feeding to youtube-dl / yt-dlp\n\nThese are all equivalent ways to achieve the same thing:\n1. Chrome: Save As | Web Page, HTML Only --\u003e doesn't work, empty page\n1. Chrome: Save As | Web page, Single File --\u003e works, full HTML, embeds images, uses \"quoted printable encoding\", i.e. `=` becomes `=3D`\n1. Chrome: Save As | Web page, Complete --\u003e works, full HTML, not encoded, saves album/playlist covers as image files.\n1. Chrome: DevTools | Elements | \u003cbody\u003e | right-click | Copy | Copy element | Paste into text editor --\u003e works, full HTML\n1. Chrome: Extensions | XPath Helper | Ctrl-Shift-X | Hover over element | Shift | Edit XPath to remove e.g. `[409]` | Append `/@href` --\u003e works, list of URLs\n1. Chrome: DevTools | Console | [\u003cload jquery\u003e](https://stackoverflow.com/a/7474386) | [\u003cXPath add-in\u003e](https://stackoverflow.com/a/20495940) | `$(document).xpathEvaluate('//body/div/foo')`\n1. Chrome: DevTools | Elements | right-click | Copy | Copy JS | (paste into console and edit - see snippet below)\n1. Chrome: Extensions | AutoHAR | chrome --auto-open-devtools-for-tabs | ...etc\n1. Chrome: DevTools | Network | Filter | Fetch/XHR | https://music.youtube.com/youtubei/v1/browse/...etc... | (a) Save all as HAR with content, (b) (down-arrow near top-right) Export HAR... \n1. (Idea) Headless chrome + puppeteer or playwright\n\nJavascript snippet:\n    \n```\nitems = document.querySelectorAll(\"#items \u003e ytmusic-two-row-item-renderer\");\nitems.forEach((item) =\u003e {\n    drill = item.querySelector(\"div.details.style-scope.ytmusic-two-row-item-renderer\");\n    span = drill.querySelector('span \u003e yt-formatted-string \u003e span:nth-child(3)');\n    if (! span) { return };\n    console.log(\n        drill.querySelector('a').toString()\n        + \"    \" + span.innerHTML\n        + \"    \" + drill.querySelector('a').text\n    );\n});\n```\n\nShorter snippet:\n```\nvar output = '';\ndocument.querySelectorAll(\"h3 \u003e div \u003e div \u003e a\").forEach((item) =\u003e { output += item.text + \"\\n\"; });\nconsole.log(output);\nconsole.save(output);\n```\n\n[Save data out of console](https://stackoverflow.com/questions/41032565/how-to-copy-the-objects-from-chrome-console-window) via clipboard or writing a file (provides `console.save()` command.\n\n### Case studies\n\n- **naive-slack-scraper**. Hypothetical code that cannot exist, as it potentially wouldn't follow terms of service. So don't look for it.\n- [pokemon-data](https://github.com/pokemon-names/pokemon-data/blob/main/data/README.md). jq examples.\n- [moar jq examples](https://wills-tech-notes.blogspot.com/2022/08/jq-cheat-sheet.html)\n\n## Discussion\n\n- If an archive of data is made, and that data cannot be viewed reasonably easily in a way similar to its original presentation by a person on the street, then it can be considered not to be viewable at all. It may as well not exist for public purposes. A possible retort is to assert \"A viewer program could be built\". But if that viewer program doesn't yet exist, then the data still can't be viewed. It's a Schroedinger's archive.\n\n## Communities\n\n- https://www.reddit.com/r/DataHoarder\n- https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-Community\n\n## Similar projects\n\n- https://github.com/iipc/awesome-web-archiving\n- https://github.com/lorien/awesome-web-scraping\n- https://github.com/igorbarinov/awesome-data-engineering\n\n## Related resources\n\n- https://pricepergig.com - Compare HDD/SSD prices by cost per TB; handy for storage-buying decisions.\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/all-the-data%2Fawesome-data-hoarding/projects"}