An open API service indexing awesome lists of open source software.

https://github.com/deutsche-nationalbibliothek/webarchiving-awesome-graph

This graph data is extracted from the Awesome Web Archiving README
https://github.com/deutsche-nationalbibliothek/webarchiving-awesome-graph

List: webarchiving-awesome-graph

dataset graph rdf webarchive

Last synced: about 1 month ago
JSON representation

This graph data is extracted from the Awesome Web Archiving README

Awesome Lists containing this project

README

          

# The Web Archive Awesome Graph (WAAG)

The original Awesome List is at: https://github.com/iipc/awesome-web-archiving

The projects can be explored here: https://aksw.github.io/awesome-store/ (very early prototype)

## Contents

- [Resources for Web Publishers](#resources-for-web-publishers)
- [Web Archiving Service Providers](#web-archiving-service-providers)
- [Self-hostable, Open Source](#self-hostable,-open-source)
- [Hosted, Closed Source](#hosted,-closed-source)
- [Training/Documentation](#training/documentation)
- [The WARC Standard](#the-warc-standard)
- [For Researchers using Web Archives](#for-researchers-using-web-archives)
- [Introductions to Web Archiving Concepts](#introductions-to-web-archiving-concepts)
- [Training Materials](#training-materials)
- [Community Resources](#community-resources)
- [Discord](#discord)
- [Mailing Lists](#mailing-lists)
- [Other Awesome Lists](#other-awesome-lists)
- [Slack](#slack)
- [Twitter](#twitter)
- [Blogs and Scholarship](#blogs-and-scholarship)
- [Public Data](#public-data)
- [Tools & Software](#tools--software)
- [Curation](#curation)
- [Replay](#replay)
- [Analysis](#analysis)
- [Quality Assurance](#quality-assurance)
- [Search & Discovery](#search--discovery)
- [WARC I/O Libraries](#warc-i/o-libraries)
- [Utilities](#utilities)
- [Acquisition](#acquisition)
- [In Development](#in-development)
- [Stable](#stable)

## Resources for Web Publishers

These resources can help when working with individuals or organisations who
publish on the web, and who want to make sure their site can be archived.
- [Definition of Web Archivability](https://nullhandle.org/web-archivability/index.html) - This describes the ease with which web content can be preserved. (

## Web Archiving Service Providers

The intention is that we only list services that allow web archives to be
exported in standard formats (WARC or WACZ). But this is not an endorsement of
these services, and readers should check and evaluate these options based on
their needs.

### Self-hostable, Open Source

- [Browsertrix](https://webrecorder.net/browsertrix/) - From
- [Conifer](https://conifer.rhizome.org/) - From

### Hosted, Closed Source

- [Archive-It](https://archive-it.org/) - From the Internet Archive.
- [Arkiwera](https://arkiwera.se/wp/websites/)
- [Hanzo](https://www.hanzo.co/chronicle)
- [MirrorWeb](https://www.mirrorweb.com/solutions/capabilities/website-archiving)
- [PageFreezer](https://www.pagefreezer.com/)
- [Smarsh](https://www.smarsh.com/platform/compliance-management/web-archive)

## Training/Documentation

This section provides a curated list of training materials, documentation, and
educational resources for those interested in learning about web archiving
practices, methodologies, and tools.

### The WARC Standard

- [Offical ISO 28500 WARC specification homepage](http://bibnum.bnf.fr/WARC/)
- [The warc-specifications](https://iipc.github.io/warc-specifications/) - A community HTML version of the official specification and hub for new proposals.

### For Researchers using Web Archives

- [Archives Unleashed Toolkit documentation](https://aut.docs.archivesunleashed.org/)
- [GLAM Workbench: Web Archives](https://glam-workbench.github.io/web-archives/) - See also
- [Tutorial for Humanities researchers about how to explore Arquivo.pt](https://sobre.arquivo.pt/en/tutorial-for-humanities-researchers-about-how-to-use-arquivo-pt/)

### Introductions to Web Archiving Concepts

- [Glossary of Archive-It and Web Archiving Terms](https://support.archive-it.org/hc/en-us/articles/208111686-Glossary-of-Archive-It-and-Web-Archiving-Terms)
- [Retrieving and Archiving Information from Websites by Wael Eskandar and Brad Murray](https://kit.exposingtheinvisible.org/en/web-archive.html/)
- [The Web Archiving Lifecycle Model](https://archive-it.org/blog/post/announcing-the-web-archiving-life-cycle-model/) - An attempt to incorporate the technological and programmatic arms of the web archiving into a framework that will be relevant to any organization seeking to archive content from the web. Archive-It, the web archiving service from the Internet Archive, developed the model based on its work with memory institutions around the world.
- [What is a web archive?](https://youtu.be/ubDHY-ynWi0) - A video from
- [Wikipedia's List of Web Archiving Initiatives](https://en.wikipedia.org/wiki/List_of_Web_archiving_initiatives)

### Training Materials

- [A Whirlwind Tour of Common Crawl's Datasets as a Python notebook](https://github.com/commoncrawl/whirlwind-python-notebook)
- [A Whirlwind Tour of Common Crawl's Datasets using Java](https://github.com/commoncrawl/whirlwind-java/)
- [A Whirlwind Tour of Common Crawl's Datasets using Python](https://github.com/commoncrawl/whirlwind-python/)
- [Continuing Education to Advance Web Archiving (CEDWARC)](https://cedwarc.github.io/)
- [IIPC and DPC Training materials: module for beginners (8 sessions)](https://netpreserve.org/web-archiving/training-materials/)
- [UNT Web Archiving Course](https://github.com/vphill/web-archiving-course)

## Community Resources

### Discord

- [Common Crawl Foundation](https://discord.gg/njaVFh7avF)

### Mailing Lists

- [Common Crawl](https://groups.google.com/g/common-crawl)
- [IIPC](http://netpreserve.org/about-us/iipc-mailing-list/)
- [OpenWayback](https://github.com/iipc/openwayback/) - The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. πŸ’½
- [OpenWayback](https://groups.google.com/g/openwayback-dev) - The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. πŸ’½
- [WASAPI](https://groups.google.com/g/wasapi-community)

### Other Awesome Lists

- [Awesome Memento](https://github.com/machawk1/awesome-memento)
- [The WARC Ecosystem](http://www.archiveteam.org/index.php?title=The_WARC_Ecosystem)
- [The Web Crawl section of COPTR](http://coptr.digipres.org/Category:Web_Crawl)
- [Web Archiving Community](https://github.com/ArchiveBox/ArchiveBox/wiki/Web-Archiving-Community)

### Slack

- [Archivers Slack](https://archivers.slack.com)
- [Archives Unleashed Slack](https://archivesunleashed.slack.com/)
- [Common Crawl Foundation Partners](https://ccfpartners.slack.com/)
- [IIPC Slack](https://iipc.slack.com/) - Ask

### Twitter

- [#WebArchiveWednesday](https://twitter.com/hashtag/webarchivewednesday)
- [#WebArchiving](https://twitter.com/search?q=%23webarchiving)
- [@NetPreserve](https://twitter.com/NetPreserve) - Official IIPC handle.
- [@WebSciDL](https://twitter.com/WebSciDL) - ODU Web Science and Digital Libraries Research Group.
- [@commoncrawl](https://twitter.com/commoncrawl) - Official Common Crawl Foundation handle.

### Blogs and Scholarship

- [Common Crawl Foundation Blog](https://commoncrawl.org/blog)
- [DSHR's Blog](https://blog.dshr.org/) - David Rosenthal regularly reviews and summarizes work done in the Digital Preservation field.
- [IIPC Blog](https://netpreserveblog.wordpress.com/)
- [The Web as History](https://www.uclpress.co.uk/products/84010) - An open-source book that provides a conceptual overview to web archiving research, as well as several case studies.
- [UK Web Archive Blog](https://blogs.bl.uk/webarchive/)
- [WS-DL Blog](https://ws-dl.blogspot.com/) - Web Science and Digital Libraries Research Group blogs about various Web archiving related topics, scholarly work, and academic trip reports.
- [Web Archiving Roundtable](https://webarchivingrt.wordpress.com/) - Unofficial blog of the Web Archiving Roundtable of the

## Public Data

This is a list of publicly available WARCs, Wayback Machines, CDX API endpoints,
other indexes, and so on.
- [Common Crawl CDX API](https://index.commoncrawl.org/)
- [Common Crawl files](https://data.commoncrawl.org/) - WARCs, CDX files, parquet url index, parquet host index, etc.
- [End of Term Archive](https://eotarchive.org/) - WARCs, CDX files, parquet url index
- [Internet Archive Wayback](https://web.archive.org/)
- [UK Government Web Archive](https://www.nationalarchives.gov.uk/webarchive/) - Wayback
- [Webrecorder US GovArchive](https://govarchive.us/) - high-fidelity replay

## Tools & Software

This list of tools and software is intended to briefly describe some of the most
important and widely-used tools related to web archiving. For more details, we
recommend you refer to (and contribute to!) these excellent resources from other
groups:
- [Awesome Website Change Monitoring](https://github.com/edgi-govdata-archiving/awesome-website-change-monitoring) πŸ’½ ⭐ 512 πŸ‘€ 25
- [Comparison of web archiving software](https://github.com/archivers-space/research/tree/master/web_archiving) πŸ’½

### Curation

- [Zotero Robust Links Extension](https://robustlinks.mementoweb.org/zotero/) - A πŸ’½

### Replay

- [InterPlanetary Wayback (ipwb)](https://github.com/oduwsdl/ipwb) - Web Archive (WARC) indexing and replay using πŸ’½ ⭐ 652 πŸ‘€ 19
- [OpenWayback](https://github.com/iipc/openwayback/) - The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. πŸ’½
- [OpenWayback](https://groups.google.com/g/openwayback-dev) - The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. πŸ’½
- [PYWB](https://github.com/webrecorder/pywb) - A Python 3 implementation of web archival replay tools, sometimes also known as 'Wayback Machine'. πŸ’½ ⭐ 1669 πŸ‘€ 58
- [Reconstructive](https://oduwsdl.github.io/Reconstructive/) - Reconstructive is a ServiceWorker module for client-side reconstruction of composite mementos by rerouting resource requests to corresponding archived copies (JavaScript). πŸ’½
- [ReplayWeb.page](https://webrecorder.net/replaywebpage/) - A browser-based, fully client-side replay engine for both local and remote WARC & WACZ files. Also available as an Electron based desktop application. πŸ’½
- [warc2html](https://github.com/iipc/warc2html) - Converts WARC files to static HTML suitable for browsing offline or rehosting. πŸ’½ ⭐ 56 πŸ‘€ 9

### Analysis

- [ArchiveSpark](https://github.com/helgeho/ArchiveSpark) - An Apache Spark framework (not only) for Web Archives that enables easy data processing, extraction as well as derivation. πŸ’½ ⭐ 161 πŸ‘€ 14
- [Archives Research Compute Hub](https://github.com/internetarchive/arch) - Web application for distributed compute analysis of Archive-It web archive collections. πŸ’½ ⭐ 20 πŸ‘€ 18
- [Archives Unleashed Notebooks](https://github.com/archivesunleashed/notebooks) - Notebooks for working with web archives with the Archives Unleashed Toolkit, and derivatives generated by the Archives Unleashed Toolkit. πŸ’½ ⭐ 26 πŸ‘€ 5
- [Archives Unleashed Toolkit](https://github.com/archivesunleashed/aut) - Archives Unleashed Toolkit (AUT) is an open-source platform for analyzing web archives with Apache Spark. πŸ’½ ⭐ 158 πŸ‘€ 12
- [Common Crawl Columnar Index](https://commoncrawl.org/tag/columnar-index/) - SQL-queryable index, with CDX info plus language classification. πŸ’½
- [Common Crawl Jupyter notebooks](https://github.com/commoncrawl/cc-notebooks) - A collection of notebooks using Common Crawl's various datasets. πŸ’½ ⭐ 66 πŸ‘€ 16
- [Common Crawl Web Graph](https://commoncrawl.org/category/web-graph/) - A host or domain-level graph of the web, with ranking information. πŸ’½
- [Tweet Archvies Unleashed Toolkit](https://github.com/archivesunleashed/twut) - An open-source toolkit for analyzing line-oriented JSON Twitter archives with Apache Spark. πŸ’½ ⭐ 10 πŸ‘€ 3
- [Web Data Commons](http://webdatacommons.org/) - Structured data extracted from Common Crawl. πŸ’½

### Quality Assurance

- [Chrome Check My Links](https://chromewebstore.google.com/detail/check-my-links/ojkcdipcgfaekbeaelaapakgnjflfglf) - Browser extension: a link checker with more options. πŸ’½
- [Chrome Open Multiple URLs](https://chromewebstore.google.com/detail/open-multiple-urls/oifijhaokejakekmnjmphonojcfkpbbh?hl=de) - Browser extension: opens multiple URLs and also extracts URLs from text. πŸ’½
- [Chrome Revolver](https://chromewebstore.google.com/detail/revolver-tabs/dlknooajieciikpedpldejhhijacnbda) - Browser extension: switches between browser tabs. πŸ’½
- [Chrome link checker](https://chromewebstore.google.com/detail/link-checker/aibjbgmpmnidnmagaefhmcjhadpffaoi) - Browser extension: basic link checker. πŸ’½
- [Chrome link gopher](https://chromewebstore.google.com/detail/bpjdkodgnbfalgghnbeggfbfjpcfamkf/publish-accepted?hl=en-US&gl=US) - Browser extension: link harvester on a page. πŸ’½
- [FlameShot](https://github.com/flameshot-org/flameshot) - Screen capture and annotation on Ubuntu. πŸ’½ ⭐ 30134 πŸ‘€ 211
- [PlayOnLinux](https://www.playonlinux.com/en/) - For running Xenu and Notepad++ on Ubuntu. πŸ’½
- [PlayOnMac](https://www.playonmac.com/en/) - For running Xenu and Notepad++ on macOS. πŸ’½
- [Windows Snipping Tool](https://support.microsoft.com/en-gb/help/13776/windows-use-snipping-tool-to-capture-screenshots) - Windows built-in for partial screen capture and annotation. On macOS you can use Command + Shift + 4 (keyboard shortcut for taking partial screen capture). πŸ’½
- [WineBottler](http://winebottler.kronenberg.org/) - For running Xenu and Notepad++ on macOS. πŸ’½
- [Xenu](http://home.snafu.de/tilman/xenulink.html) - Desktop link checker for Windows. πŸ’½
- [xDoTool](https://github.com/jordansissel/xdotool) - Click automation on Ubuntu. πŸ’½ ⭐ 3813 πŸ‘€ 62

### Search & Discovery

- [Mink](https://github.com/machawk1/Mink) - A πŸ’½ ⭐ 58 πŸ‘€ 3
- [PANDORΓ†](https://github.com/Guillaume-Levrier/PANDORAE) - A desktop research software to be plugged on a Solr endpoint to query, retrieve, normalize and visually explore web archives. πŸ’½ ⭐ 16 πŸ‘€ 1
- [SecurityTrails](https://securitytrails.com/) - Web based archive for WHOIS and DNS records. REST API available free of charge. πŸ’½
- [Shine](https://github.com/ukwa/shine) - A prototype web archives exploration UI, developed with researchers as part of the πŸ’½ ⭐ 43 πŸ‘€ 2
- [SolrWayback](https://github.com/netarchivesuite/solrwayback) - A backend Java and frontend VUE JS project with freetext search and a build in playback engine. Require Warc files has been index with the Warc-Indexer. The web application also has a wide range of data visualization tools and data export tools that can be used on the whole webarchive. πŸ’½ ⭐ 144 πŸ‘€ 20
- [Tempas v1](http://tempas.L3S.de/v1) - Temporal web archive search based on πŸ’½
- [Tempas v2](http://tempas.L3S.de/v2) - Temporal web archive search based on links and anchor texts extracted from the German web from 1996 to 2013 (results are not limited to German pages, e.g., πŸ’½
- [Warclight](https://github.com/archivesunleashed/warclight) - A Project Blacklight based Rails engine that supports the discovery of web archives held in the WARC and ARC formats. πŸ’½ ⭐ 50 πŸ‘€ 3
- [Wasp](https://github.com/webis-de/wasp) - A fully functional prototype of a personal πŸ’½ ⭐ 28 πŸ‘€ 10
- [hyphe](https://github.com/medialab/hyphe) - A webcrawler built for research uses with a graphical user interface in order to build web corpuses made of lists of web actors and maps of links between them. πŸ’½ ⭐ 381 πŸ‘€ 28
- [playback](https://github.com/wabarc/playback) - A toolkit for searching archived webpages from πŸ’½ ⭐ 13 πŸ‘€ 2
- [webarchive-discovery](https://github.com/ukwa/webarchive-discovery) - WARC and ARC full-text indexing and discovery tools, with a number of associated tools capable of using the index shown below. πŸ’½ ⭐ 132 πŸ‘€ 20

### WARC I/O Libraries

- [FastWARC](https://github.com/chatnoir-eu/chatnoir-resiliparse) - A high-performance WARC parsing library (Python). πŸ’½ ⭐ 141 πŸ‘€ 9
- [HadoopConcatGz](https://github.com/helgeho/HadoopConcatGz) - A Splitable Hadoop InputFormat for Concatenated GZIP Files (and πŸ’½ ⭐ 9 πŸ‘€ 1
- [Jwat](https://github.com/netarchivesuite/jwat) - Libraries for reading/writing/validating WARC/ARC/GZIP files (Java). πŸ’½ ⭐ 4 πŸ‘€ 6
- [Jwat-Tools](https://github.com/netarchivesuite/jwat-tools) - Tools for reading/writing/validating WARC/ARC/GZIP files (Java). πŸ’½ ⭐ 5 πŸ‘€ 5
- [Sparkling](https://github.com/internetarchive/Sparkling) - Internet Archive's Sparkling Data Processing Library. πŸ’½ ⭐ 17 πŸ‘€ 18
- [Unwarcit](https://github.com/emmadickson/unwarcit) - Command line interface to unzip WARC and WACZ files (Python). πŸ’½ ⭐ 13 πŸ‘€ 4
- [Warcat](https://github.com/chfoo/warcat) - Tool and library for handling Web ARChive (WARC) files (Python). πŸ’½ ⭐ 165 πŸ‘€ 9
- [Warcat-rs](https://github.com/chfoo/warcat-rs) - Command-line tool and Rust library for handling Web ARChive (WARC) files. πŸ’½ ⭐ 31 πŸ‘€ 1
- [jwarc](https://github.com/iipc/jwarc) - Read and write WARC files with a type safe API (Java). πŸ’½ ⭐ 59 πŸ‘€ 4
- [node-warc](https://github.com/N0taN3rd/node-warc) - Parse WARC files or create WARC files using either πŸ’½ ⭐ 104 πŸ‘€ 7
- [warc](https://github.com/jedireza/warc) - A Rust library for reading and writing WARC files. πŸ’½ ⭐ 59 πŸ‘€ 2
- [warcio](https://github.com/webrecorder/warcio) - Streaming WARC/ARC library for fast web archive IO (Python). πŸ’½ ⭐ 458 πŸ‘€ 21
- [warctools](https://github.com/internetarchive/warctools) - Library to work with ARC and WARC files (Python). πŸ’½ ⭐ 174 πŸ‘€ 38
- [webarchive](https://github.com/richardlehane/webarchive) - Golang readers for ARC and WARC webarchive formats (Golang). πŸ’½ ⭐ 20 πŸ‘€ 6

### Utilities

- [ArchiveTools](https://github.com/recrm/ArchiveTools) - Collection of tools to extract and interact with WARC files (Python). πŸ’½ ⭐ 78 πŸ‘€ 5
- [Go Get Crawl](https://github.com/karust/gogetcrawl) - Extract web archive data using πŸ’½ ⭐ 180 πŸ‘€ 2
- [HTTPreserve linkstat](https://github.com/httpreserve/linkstat) - Command line implementation of πŸ’½ ⭐ 10 πŸ‘€ 1
- [Internet Archive Library](https://github.com/jjjake/internetarchive) - A command line tool and Python library for interacting directly with πŸ’½ ⭐ 1873 πŸ‘€ 53
- [MemGator](https://github.com/oduwsdl/MemGator) - A Memento Aggregator CLI and Server (Golang). πŸ’½ ⭐ 80 πŸ‘€ 9
- [MementoMap](https://github.com/oduwsdl/MementoMap) - A Tool to Summarize Web Archive Holdings (Python). πŸ’½ ⭐ 11 πŸ‘€ 5
- [OutbackCDX](https://github.com/nla/outbackcdx) - RocksDB-based capture index (CDX) server supporting incremental updates and compression. Can be used as backend for OpenWayback, PyWb and πŸ’½ ⭐ 43 πŸ‘€ 19
- [The Unarchiver](https://theunarchiver.com/) - Program to extract the contents of many archive formats, inclusive of WARC, to a file system. Free variant of The Archive Browser (macOS only, Proprietary app). πŸ’½
- [WarcPartitioner](https://github.com/helgeho/WarcPartitioner) - Partition (W)ARC Files by MIME Type and Year. πŸ’½ ⭐ 1 πŸ‘€ 1
- [Warchaeology](https://nlnwa.github.io/warchaeology/) - Warchaeology is a collection of tools for inspecting, manipulating, deduplicating and validating WARC-files. πŸ’½
- [bagnabit2warc](https://github.com/internetarchive/bagnabit2warc) - Convert a πŸ’½ ⭐ 1
- [cdx-toolkit](https://pypi.org/project/cdx-toolkit/) - Library and CLI to consult cdx indexes and create WARC extractions of subsets. Abstracts away Common Crawl's unusual crawl structure. πŸ’½
- [duckdb-web-archive-cdx](https://github.com/midwork-finds-jobs/duckdb-web-archive) - DuckDB extension to query the Internet Archive and CommonCrawl CDX APIs directly from SQL. πŸ’½ ⭐ 21
- [duckdb_warc](https://github.com/midwork-finds-jobs/duckdb_warc) - DuckDB extension to query WARC files. πŸ’½ ⭐ 4
- [gowarcserver](https://github.com/nlnwa/gowarcserver) πŸ’½ ⭐ 17 πŸ‘€ 5
- [har2warc](https://github.com/webrecorder/har2warc) - Convert HTTP Archive (HAR) -> Web Archive (WARC) format (Python). πŸ’½ ⭐ 55 πŸ‘€ 6
- [httpreserve.info](https://httpreserve.info) - Service to return the status of a web page or save it to the Internet Archive. HTTPreserve includes disambiguation of well-known short link services. It returns JSON via the browser or command line via CURL using GET. Describes web sites using earliest and latest dates in the Internet Archive and demonstrates the construction of Robust Links in its output using that range. (Golang). πŸ’½
- [httrack2warc](https://github.com/nla/httrack2warc) - Convert HTTrack archives to WARC format (Java). πŸ’½ ⭐ 34 πŸ‘€ 18
- [node-cdxj](https://github.com/N0taN3rd/node-cdxj) πŸ’½ ⭐ 2 πŸ‘€ 2
- [py-wasapi-client](https://github.com/unt-libraries/py-wasapi-client) - Command line application to download crawls from WASAPI (Python). πŸ’½ ⭐ 16 πŸ‘€ 4
- [tikalinkextract](https://github.com/httpreserve/tikalinkextract) - Extract hyperlinks as a seed for web archiving from folders of document types that can be parsed by Apache Tika (Golang, Apache Tika Server). πŸ’½ ⭐ 11 πŸ‘€ 2
- [warc-safe](https://github.com/natliblux/warc-safe) - Automatic detection of viruses and NSFW content in WARC files. πŸ’½ ⭐ 18 πŸ‘€ 4
- [warcbench](https://github.com/harvard-lil/warcbench) - A tool for exploring, analyzing, transforming, recombining, and extracting data from WARC (Web ARChive) files. πŸ’½ ⭐ 19 πŸ‘€ 2
- [warcdb](https://github.com/florents-Tselai/warcdb) - A command line utility (Python) for importing WARC files into a SQLite database. πŸ’½ ➑️ moved to https://github.com/Florents-Tselai/WarcDB
- [warcdedupe](https://gitlab.com/taricorp/warcdedupe) - WARC deduplication tool (and WARC library) written in Rust. πŸ’½
- [warcrefs](https://github.com/arcalex/warcrefs) - Web archive deduplication tools. πŸ’½ ⭐ 10 πŸ‘€ 4
- [wasapi-downloader](https://github.com/sul-dlss/wasapi-downloader) - Java command line application to download crawls from WASAPI. πŸ’½ ➑️ moved to https://github.com/sul-dlss-deprecated/wasapi-downloader
- [webarchive-indexing](https://github.com/ikreymer/webarchive-indexing) - Tools for bulk indexing of WARC/ARC files on Hadoop, EMR or local file system. πŸ’½ ⭐ 46 πŸ‘€ 8
- [wikiteam](https://github.com/WikiTeam/wikiteam) - Tools for downloading and preserving wikis. πŸ’½ ⭐ 848 πŸ‘€ 38

### Acquisition

- [ArchiveBox](https://github.com/ArchiveBox/ArchiveBox) - A tool which maintains an additive archive from RSS feeds, bookmarks, and links using wget, Chrome headless, and other methods (formerly πŸ’½ ⭐ 27693 πŸ‘€ 181
- [ArchiveWeb.Page](https://webrecorder.net/archivewebpage/) - A plugin for Chrome and other Chromium based browsers that lets you interactively archive web pages, replay them, and export them as WARC & WACZ files. Also available as an Electron based desktop application. πŸ’½
- [Auto Archiver](https://github.com/bellingcat/auto-archiver) - Python script to automatically archive social media posts, videos, and images from a Google Sheets document. Read the πŸ’½ ⭐ 1086 πŸ‘€ 23
- [Browsertrix Crawler](https://github.com/webrecorder/browsertrix-crawler) - A Chromium based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. πŸ’½ ⭐ 1053 πŸ‘€ 23
- [Brozzler](https://github.com/internetarchive/brozzler) - A distributed web crawler (ηˆ¬θ™«) that uses a real browser (Chrome or Chromium) to fetch pages and embedded urls and to extract links. πŸ’½ ⭐ 799 πŸ‘€ 34
- [Cairn](https://github.com/wabarc/cairn) - A npm package and CLI tool for saving webpages. πŸ’½ ⭐ 52 πŸ‘€ 2
- [Chronicler](https://github.com/CGamesPlay/chronicler) - Web browser with record and replay functionality. πŸ’½ ⭐ 92 πŸ‘€ 1
- [Community Archive](https://www.community-archive.org/) - Open Twitter Database and API with tools and resources for building on archived Twitter data. πŸ’½
- [Crawl](https://git.autistici.org/ale/crawl) - A simple web crawler in Golang. πŸ’½
- [DiskerNet](https://github.com/DO-SAY-GO/dn) - A non-WARC-based tool which hooks into the Chrome browser and archives everything you browse making it available for offline replay. πŸ’½ ⭐ 3903 πŸ‘€ 37
- [F(b)arc](https://github.com/justinlittman/fbarc) - A commandline tool and Python library for archiving data from πŸ’½ ⭐ 78 πŸ‘€ 1
- [HTTrack](http://www.httrack.com/) - An open source website copying utility. πŸ’½
- [Heritrix](https://github.com/internetarchive/heritrix3/wiki) - An open source, extensible, web-scale, archival quality web crawler. πŸ’½
- [Obelisk](https://github.com/go-shiori/obelisk) - Go package and CLI tool for saving web page as single HTML file. πŸ’½ ⭐ 315 πŸ‘€ 8
- [Scoop](https://github.com/harvard-lil/scoop) - High-fidelity, browser-based, single-page web archiving library and CLI for witnessing the web. πŸ’½ ⭐ 200 πŸ‘€ 7
- [SingleFile](https://github.com/gildas-lormeau/SingleFile) - Browser extension for Firefox/Chrome and CLI tool to save a faithful copy of a complete page as a single HTML file. πŸ’½ ⭐ 21491 πŸ‘€ 139
- [SiteStory](http://mementoweb.github.io/SiteStory/) - A transactional archive that selectively captures and stores transactions that take place between a web client (browser) and a web server. πŸ’½
- [Social Feed Manager](https://gwu-libraries.github.io/sfm-ui/) - Open source software that enables users to create social media collections from Twitter, Tumblr, Flickr, and Sina Weibo public APIs. πŸ’½
- [Squidwarc](https://github.com/N0taN3rd/Squidwarc) - An πŸ’½ ⭐ 175 πŸ‘€ 9
- [StormCrawler](http://stormcrawler.net/) - A collection of resources for building low-latency, scalable web crawlers on Apache Storm. πŸ’½
- [WAIL](https://github.com/machawk1/wail) - A graphical user interface (GUI) atop multiple web archiving tools intended to be used as an easy way for anyone to preserve and replay web pages; πŸ’½ ⭐ 397 πŸ‘€ 10
- [WARCreate](http://matkelly.com/warcreate/) - A πŸ’½
- [Warcprox](https://github.com/internetarchive/warcprox) - WARC-writing MITM HTTP/S proxy. πŸ’½ ⭐ 453 πŸ‘€ 30
- [Warcworker](https://github.com/peterk/warcworker) - An open source, dockerized, queued, high fidelity web archiver based on Squidwarc with a simple web GUI. πŸ’½ ⭐ 62 πŸ‘€ 5
- [Wayback](https://github.com/wabarc/wayback) - A toolkit for snapshot webpage to Internet Archive, archive.today, IPFS and beyond. πŸ’½ ⭐ 2200 πŸ‘€ 6
- [Waybackpy](https://github.com/akamhy/waybackpy) - Wayback Machine Save, CDX and availability API interface in Python and a command-line tool πŸ’½ ⭐ 587 πŸ‘€ 8
- [Web2Warc](https://github.com/helgeho/Web2Warc) - An easy-to-use and highly customizable crawler that enables anyone to create their own little Web archives (WARC/CDX). πŸ’½ ⭐ 26 πŸ‘€ 2
- [WebMemex](https://github.com/WebMemex) - Browser extension for Firefox and Chrome which lets you archive web pages you visit. πŸ’½
- [Web Curator Tool](https://webcuratortool.org) - Open-source workflow management for selective web archiving. πŸ’½
- [Wget](http://www.gnu.org/software/wget/) - An open source file retrieval utility that of πŸ’½
- [Wget-lua](https://github.com/alard/wget-lua) - Wget with Lua extension. πŸ’½ ⭐ 24 πŸ‘€ 3
- [Wpull](https://github.com/ArchiveTeam/wpull) - A Wget-compatible (or remake/clone/replacement/alternative) web downloader and crawler. πŸ’½ ⭐ 610 πŸ‘€ 21
- [archivenow](https://github.com/oduwsdl/archivenow) - A πŸ’½ ⭐ 434 πŸ‘€ 19
- [crau](https://github.com/turicas/crau) - crau is the way (most) Brazilians pronounce crawl, it's the easiest command-line tool for archiving the Web and playing archives: you just need a list of URLs. πŸ’½ ⭐ 64 πŸ‘€ 2
- [crocoite](https://github.com/PromyLOPh/crocoite) - Crawl websites using headless Google Chrome/Chromium and save resources, static DOM snapshot and page screenshots to WARC files. πŸ’½ ⭐ 45 πŸ‘€ 1
- [freeze-dry](https://github.com/WebMemex/freeze-dry) - JavaScript library to turn page into static, self-contained HTML document; useful for browser extensions. πŸ’½ ⭐ 301 πŸ‘€ 10
- [grab-site](https://github.com/ArchiveTeam/grab-site) - The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns. πŸ’½ ⭐ 1586 πŸ‘€ 39
- [html2warc](https://github.com/steffenfritz/html2warc) - A simple script to convert offline data into a single WARC file. πŸ’½ ⭐ 24 πŸ‘€ 2
- [monolith](https://github.com/Y2Z/monolith) - CLI tool to save a web page as a single HTML file. πŸ’½ ⭐ 15188 πŸ‘€ 64
- [twarc](https://github.com/DocNow/twarc) - A command line tool and Python library for archiving Twitter JSON data. πŸ’½ ⭐ 1394 πŸ‘€ 33

## In Development

- [ArchiveBox](https://github.com/ArchiveBox/ArchiveBox) - A tool which maintains an additive archive from RSS feeds, bookmarks, and links using wget, Chrome headless, and other methods (formerly πŸ’½ ⭐ 27693 πŸ‘€ 181
- [Chronicler](https://github.com/CGamesPlay/chronicler) - Web browser with record and replay functionality. πŸ’½ ⭐ 92 πŸ‘€ 1
- [DiskerNet](https://github.com/DO-SAY-GO/dn) - A non-WARC-based tool which hooks into the Chrome browser and archives everything you browse making it available for offline replay. πŸ’½ ⭐ 3903 πŸ‘€ 37
- [MementoMap](https://github.com/oduwsdl/MementoMap) - A Tool to Summarize Web Archive Holdings (Python). πŸ’½ ⭐ 11 πŸ‘€ 5
- [Squidwarc](https://github.com/N0taN3rd/Squidwarc) - An πŸ’½ ⭐ 175 πŸ‘€ 9
- [Tweet Archvies Unleashed Toolkit](https://github.com/archivesunleashed/twut) - An open-source toolkit for analyzing line-oriented JSON Twitter archives with Apache Spark. πŸ’½ ⭐ 10 πŸ‘€ 3
- [Warcat-rs](https://github.com/chfoo/warcat-rs) - Command-line tool and Rust library for handling Web ARChive (WARC) files. πŸ’½ ⭐ 31 πŸ‘€ 1
- [Warclight](https://github.com/archivesunleashed/warclight) - A Project Blacklight based Rails engine that supports the discovery of web archives held in the WARC and ARC formats. πŸ’½ ⭐ 50 πŸ‘€ 3
- [Wasp](https://github.com/webis-de/wasp) - A fully functional prototype of a personal πŸ’½ ⭐ 28 πŸ‘€ 10
- [WebMemex](https://github.com/WebMemex) - Browser extension for Firefox and Chrome which lets you archive web pages you visit. πŸ’½
- [crocoite](https://github.com/PromyLOPh/crocoite) - Crawl websites using headless Google Chrome/Chromium and save resources, static DOM snapshot and page screenshots to WARC files. πŸ’½ ⭐ 45 πŸ‘€ 1
- [duckdb-web-archive-cdx](https://github.com/midwork-finds-jobs/duckdb-web-archive) - DuckDB extension to query the Internet Archive and CommonCrawl CDX APIs directly from SQL. πŸ’½ ⭐ 21
- [duckdb_warc](https://github.com/midwork-finds-jobs/duckdb_warc) - DuckDB extension to query WARC files. πŸ’½ ⭐ 4
- [freeze-dry](https://github.com/WebMemex/freeze-dry) - JavaScript library to turn page into static, self-contained HTML document; useful for browser extensions. πŸ’½ ⭐ 301 πŸ‘€ 10
- [playback](https://github.com/wabarc/playback) - A toolkit for searching archived webpages from πŸ’½ ⭐ 13 πŸ‘€ 2
- [tikalinkextract](https://github.com/httpreserve/tikalinkextract) - Extract hyperlinks as a seed for web archiving from folders of document types that can be parsed by Apache Tika (Golang, Apache Tika Server). πŸ’½ ⭐ 11 πŸ‘€ 2
- [warcdedupe](https://gitlab.com/taricorp/warcdedupe) - WARC deduplication tool (and WARC library) written in Rust. πŸ’½

## Stable

- [ArchiveSpark](https://github.com/helgeho/ArchiveSpark) - An Apache Spark framework (not only) for Web Archives that enables easy data processing, extraction as well as derivation. πŸ’½ ⭐ 161 πŸ‘€ 14
- [Archives Research Compute Hub](https://github.com/internetarchive/arch) - Web application for distributed compute analysis of Archive-It web archive collections. πŸ’½ ⭐ 20 πŸ‘€ 18
- [Archives Unleashed Notebooks](https://github.com/archivesunleashed/notebooks) - Notebooks for working with web archives with the Archives Unleashed Toolkit, and derivatives generated by the Archives Unleashed Toolkit. πŸ’½ ⭐ 26 πŸ‘€ 5
- [Archives Unleashed Toolkit](https://github.com/archivesunleashed/aut) - Archives Unleashed Toolkit (AUT) is an open-source platform for analyzing web archives with Apache Spark. πŸ’½ ⭐ 158 πŸ‘€ 12
- [Browsertrix Crawler](https://github.com/webrecorder/browsertrix-crawler) - A Chromium based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. πŸ’½ ⭐ 1053 πŸ‘€ 23
- [Brozzler](https://github.com/internetarchive/brozzler) - A distributed web crawler (ηˆ¬θ™«) that uses a real browser (Chrome or Chromium) to fetch pages and embedded urls and to extract links. πŸ’½ ⭐ 799 πŸ‘€ 34
- [Cairn](https://github.com/wabarc/cairn) - A npm package and CLI tool for saving webpages. πŸ’½ ⭐ 52 πŸ‘€ 2
- [Common Crawl Columnar Index](https://commoncrawl.org/tag/columnar-index/) - SQL-queryable index, with CDX info plus language classification. πŸ’½
- [Common Crawl Jupyter notebooks](https://github.com/commoncrawl/cc-notebooks) - A collection of notebooks using Common Crawl's various datasets. πŸ’½ ⭐ 66 πŸ‘€ 16
- [Common Crawl Web Graph](https://commoncrawl.org/category/web-graph/) - A host or domain-level graph of the web, with ranking information. πŸ’½
- [Crawl](https://git.autistici.org/ale/crawl) - A simple web crawler in Golang. πŸ’½
- [F(b)arc](https://github.com/justinlittman/fbarc) - A commandline tool and Python library for archiving data from πŸ’½ ⭐ 78 πŸ‘€ 1
- [Go Get Crawl](https://github.com/karust/gogetcrawl) - Extract web archive data using πŸ’½ ⭐ 180 πŸ‘€ 2
- [HTTPreserve linkstat](https://github.com/httpreserve/linkstat) - Command line implementation of πŸ’½ ⭐ 10 πŸ‘€ 1
- [HTTrack](http://www.httrack.com/) - An open source website copying utility. πŸ’½
- [HadoopConcatGz](https://github.com/helgeho/HadoopConcatGz) - A Splitable Hadoop InputFormat for Concatenated GZIP Files (and πŸ’½ ⭐ 9 πŸ‘€ 1
- [Heritrix](https://github.com/internetarchive/heritrix3/wiki) - An open source, extensible, web-scale, archival quality web crawler. πŸ’½
- [Internet Archive Library](https://github.com/jjjake/internetarchive) - A command line tool and Python library for interacting directly with πŸ’½ ⭐ 1873 πŸ‘€ 53
- [Jwat](https://github.com/netarchivesuite/jwat) - Libraries for reading/writing/validating WARC/ARC/GZIP files (Java). πŸ’½ ⭐ 4 πŸ‘€ 6
- [Jwat-Tools](https://github.com/netarchivesuite/jwat-tools) - Tools for reading/writing/validating WARC/ARC/GZIP files (Java). πŸ’½ ⭐ 5 πŸ‘€ 5
- [MemGator](https://github.com/oduwsdl/MemGator) - A Memento Aggregator CLI and Server (Golang). πŸ’½ ⭐ 80 πŸ‘€ 9
- [Mink](https://github.com/machawk1/Mink) - A πŸ’½ ⭐ 58 πŸ‘€ 3
- [Obelisk](https://github.com/go-shiori/obelisk) - Go package and CLI tool for saving web page as single HTML file. πŸ’½ ⭐ 315 πŸ‘€ 8
- [OpenWayback](https://github.com/iipc/openwayback/) - The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. πŸ’½
- [OpenWayback](https://groups.google.com/g/openwayback-dev) - The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. πŸ’½
- [OutbackCDX](https://github.com/nla/outbackcdx) - RocksDB-based capture index (CDX) server supporting incremental updates and compression. Can be used as backend for OpenWayback, PyWb and πŸ’½ ⭐ 43 πŸ‘€ 19
- [PANDORΓ†](https://github.com/Guillaume-Levrier/PANDORAE) - A desktop research software to be plugged on a Solr endpoint to query, retrieve, normalize and visually explore web archives. πŸ’½ ⭐ 16 πŸ‘€ 1
- [PYWB](https://github.com/webrecorder/pywb) - A Python 3 implementation of web archival replay tools, sometimes also known as 'Wayback Machine'. πŸ’½ ⭐ 1669 πŸ‘€ 58
- [ReplayWeb.page](https://webrecorder.net/replaywebpage/) - A browser-based, fully client-side replay engine for both local and remote WARC & WACZ files. Also available as an Electron based desktop application. πŸ’½
- [Scoop](https://github.com/harvard-lil/scoop) - High-fidelity, browser-based, single-page web archiving library and CLI for witnessing the web. πŸ’½ ⭐ 200 πŸ‘€ 7
- [Shine](https://github.com/ukwa/shine) - A prototype web archives exploration UI, developed with researchers as part of the πŸ’½ ⭐ 43 πŸ‘€ 2
- [SingleFile](https://github.com/gildas-lormeau/SingleFile) - Browser extension for Firefox/Chrome and CLI tool to save a faithful copy of a complete page as a single HTML file. πŸ’½ ⭐ 21491 πŸ‘€ 139
- [SiteStory](http://mementoweb.github.io/SiteStory/) - A transactional archive that selectively captures and stores transactions that take place between a web client (browser) and a web server. πŸ’½
- [Social Feed Manager](https://gwu-libraries.github.io/sfm-ui/) - Open source software that enables users to create social media collections from Twitter, Tumblr, Flickr, and Sina Weibo public APIs. πŸ’½
- [Sparkling](https://github.com/internetarchive/Sparkling) - Internet Archive's Sparkling Data Processing Library. πŸ’½ ⭐ 17 πŸ‘€ 18
- [StormCrawler](http://stormcrawler.net/) - A collection of resources for building low-latency, scalable web crawlers on Apache Storm. πŸ’½
- [Tempas v1](http://tempas.L3S.de/v1) - Temporal web archive search based on πŸ’½
- [Tempas v2](http://tempas.L3S.de/v2) - Temporal web archive search based on links and anchor texts extracted from the German web from 1996 to 2013 (results are not limited to German pages, e.g., πŸ’½
- [WAIL](https://github.com/machawk1/wail) - A graphical user interface (GUI) atop multiple web archiving tools intended to be used as an easy way for anyone to preserve and replay web pages; πŸ’½ ⭐ 397 πŸ‘€ 10
- [WARCreate](http://matkelly.com/warcreate/) - A πŸ’½
- [WarcPartitioner](https://github.com/helgeho/WarcPartitioner) - Partition (W)ARC Files by MIME Type and Year. πŸ’½ ⭐ 1 πŸ‘€ 1
- [Warcat](https://github.com/chfoo/warcat) - Tool and library for handling Web ARChive (WARC) files (Python). πŸ’½ ⭐ 165 πŸ‘€ 9
- [Warchaeology](https://nlnwa.github.io/warchaeology/) - Warchaeology is a collection of tools for inspecting, manipulating, deduplicating and validating WARC-files. πŸ’½
- [Warcprox](https://github.com/internetarchive/warcprox) - WARC-writing MITM HTTP/S proxy. πŸ’½ ⭐ 453 πŸ‘€ 30
- [Warcworker](https://github.com/peterk/warcworker) - An open source, dockerized, queued, high fidelity web archiver based on Squidwarc with a simple web GUI. πŸ’½ ⭐ 62 πŸ‘€ 5
- [Wayback](https://github.com/wabarc/wayback) - A toolkit for snapshot webpage to Internet Archive, archive.today, IPFS and beyond. πŸ’½ ⭐ 2200 πŸ‘€ 6
- [Waybackpy](https://github.com/akamhy/waybackpy) - Wayback Machine Save, CDX and availability API interface in Python and a command-line tool πŸ’½ ⭐ 587 πŸ‘€ 8
- [Web2Warc](https://github.com/helgeho/Web2Warc) - An easy-to-use and highly customizable crawler that enables anyone to create their own little Web archives (WARC/CDX). πŸ’½ ⭐ 26 πŸ‘€ 2
- [Web Curator Tool](https://webcuratortool.org) - Open-source workflow management for selective web archiving. πŸ’½
- [Web Data Commons](http://webdatacommons.org/) - Structured data extracted from Common Crawl. πŸ’½
- [Wget](http://www.gnu.org/software/wget/) - An open source file retrieval utility that of πŸ’½
- [Wget-lua](https://github.com/alard/wget-lua) - Wget with Lua extension. πŸ’½ ⭐ 24 πŸ‘€ 3
- [Wpull](https://github.com/ArchiveTeam/wpull) - A Wget-compatible (or remake/clone/replacement/alternative) web downloader and crawler. πŸ’½ ⭐ 610 πŸ‘€ 21
- [archivenow](https://github.com/oduwsdl/archivenow) - A πŸ’½ ⭐ 434 πŸ‘€ 19
- [cdx-toolkit](https://pypi.org/project/cdx-toolkit/) - Library and CLI to consult cdx indexes and create WARC extractions of subsets. Abstracts away Common Crawl's unusual crawl structure. πŸ’½
- [crau](https://github.com/turicas/crau) - crau is the way (most) Brazilians pronounce crawl, it's the easiest command-line tool for archiving the Web and playing archives: you just need a list of URLs. πŸ’½ ⭐ 64 πŸ‘€ 2
- [grab-site](https://github.com/ArchiveTeam/grab-site) - The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns. πŸ’½ ⭐ 1586 πŸ‘€ 39
- [html2warc](https://github.com/steffenfritz/html2warc) - A simple script to convert offline data into a single WARC file. πŸ’½ ⭐ 24 πŸ‘€ 2
- [httpreserve.info](https://httpreserve.info) - Service to return the status of a web page or save it to the Internet Archive. HTTPreserve includes disambiguation of well-known short link services. It returns JSON via the browser or command line via CURL using GET. Describes web sites using earliest and latest dates in the Internet Archive and demonstrates the construction of Robust Links in its output using that range. (Golang). πŸ’½
- [hyphe](https://github.com/medialab/hyphe) - A webcrawler built for research uses with a graphical user interface in order to build web corpuses made of lists of web actors and maps of links between them. πŸ’½ ⭐ 381 πŸ‘€ 28
- [monolith](https://github.com/Y2Z/monolith) - CLI tool to save a web page as a single HTML file. πŸ’½ ⭐ 15188 πŸ‘€ 64
- [node-cdxj](https://github.com/N0taN3rd/node-cdxj) πŸ’½ ⭐ 2 πŸ‘€ 2
- [node-warc](https://github.com/N0taN3rd/node-warc) - Parse WARC files or create WARC files using either πŸ’½ ⭐ 104 πŸ‘€ 7
- [py-wasapi-client](https://github.com/unt-libraries/py-wasapi-client) - Command line application to download crawls from WASAPI (Python). πŸ’½ ⭐ 16 πŸ‘€ 4
- [twarc](https://github.com/DocNow/twarc) - A command line tool and Python library for archiving Twitter JSON data. πŸ’½ ⭐ 1394 πŸ‘€ 33
- [warc](https://github.com/jedireza/warc) - A Rust library for reading and writing WARC files. πŸ’½ ⭐ 59 πŸ‘€ 2
- [warcdb](https://github.com/florents-Tselai/warcdb) - A command line utility (Python) for importing WARC files into a SQLite database. πŸ’½ ➑️ moved to https://github.com/Florents-Tselai/WarcDB
- [warcio](https://github.com/webrecorder/warcio) - Streaming WARC/ARC library for fast web archive IO (Python). πŸ’½ ⭐ 458 πŸ‘€ 21
- [warcrefs](https://github.com/arcalex/warcrefs) - Web archive deduplication tools. πŸ’½ ⭐ 10 πŸ‘€ 4
- [wasapi-downloader](https://github.com/sul-dlss/wasapi-downloader) - Java command line application to download crawls from WASAPI. πŸ’½ ➑️ moved to https://github.com/sul-dlss-deprecated/wasapi-downloader
- [webarchive-discovery](https://github.com/ukwa/webarchive-discovery) - WARC and ARC full-text indexing and discovery tools, with a number of associated tools capable of using the index shown below. πŸ’½ ⭐ 132 πŸ‘€ 20
- [wikiteam](https://github.com/WikiTeam/wikiteam) - Tools for downloading and preserving wikis. πŸ’½ ⭐ 848 πŸ‘€ 38