Wayback-Archive¶
Wayback-Archive is a command-line tool that rescues a whole website from the Wayback Machine. Give it one archived URL and you get a folder that opens in a browser and looks like the site did on that day: every page and asset the crawl can reach, links rewritten to the local copies, the Wayback Machine's toolbar and URL prefixes gone. It installs from PyPI in one command (Getting started) and every setting is an environment variable (Configuration).
-
pipx install wayback-archive, or pip, or a checkout. Python 3.10 or newer. -
One URL, one command, a quick test with
MAX_FILES, and how to open the copy on macOS, Linux and Windows. -
Every environment variable and its default: output folder, minification, trackers, links, www or not.
-
The crawl steps, the nearby-timestamp fallback and what the tool talks to.
The result¶
A snapshot of python.org from 31 December 2005, rescued with the quick-start command (MAX_FILES=50) and opened from the copy. No Wayback toolbar, no web.archive.org in the links, the stylesheet and images served from the folder on disk.





What it saves¶
- Every page and asset the crawl can reach from the URL you give it: HTML, CSS, JavaScript, images and fonts. Links are followed in HTML, in CSS (
@import,url()) and in JavaScript. - Links rewritten to relative paths, so pages open from the folder and lead to each other without touching the Wayback Machine again.
- Pages and files the snapshot is missing, when one of up to three other Wayback timestamps (a day either side and a week earlier) has them. A capture that is an archived error or a Cloudflare challenge page is replaced with the nearest good capture.
- Google Fonts, saved locally so text renders offline. A font file the archive returns as an HTML error page is dropped instead of breaking the stylesheet.
- jQuery, when the archive lost it, fetched from
code.jquery.comso the page still works. - Social icon groups, button links and cookie banners, kept working in the copy.
- URLs held in
data-*attributes, such as lazy-loaded images and videos.
Analytics and ads are removed by default (REMOVE_TRACKERS, REMOVE_ADS); external iframes only when REMOVE_EXTERNAL_IFRAMES=true. Optional minification of HTML, CSS, JavaScript and images makes the folder smaller. The full list is on Features, with the comparison against wget --mirror and httrack.
How it runs¶
flowchart LR
WB[Wayback Machine<br/>web.archive.org]
CLI[wayback-archive<br/>one process, settings from the environment]
Q[Queue of discovered URLs]
OUT[(OUTPUT_DIR<br/>./output by default)]
FB[Fallbacks:<br/>other Wayback timestamps,<br/>the nearest good capture,<br/>Google Fonts, code.jquery.com,<br/>the Squarespace CDN]
WB --> CLI
CLI --> Q --> CLI
CLI --> OUT
CLI -. only when a file is missing .-> FB
- The page at
WAYBACK_URLis fetched first. - Its HTML, CSS and JavaScript are parsed for links and every new URL joins the queue.
- Each queued file is downloaded, parsed the same way, rewritten and written under
OUTPUT_DIR. - A 404 from the archive is retried at up to three other timestamps. A file still missing is fetched live only when it is hosted on Google Fonts,
code.jquery.comor the Squarespace CDN; anything else is counted as failed. - When the queue is empty the run prints a
Download Complete!block with the files downloaded, failed and skipped.
MAX_FILES stops the crawl after trying that many files, failed ones included, which is the way to try a big site without waiting for all of it. The crawl fetches files in the order it finds them, so on a site whose stylesheet is only reached through @import a short run can stop before it; the quick start's snapshot links its stylesheet directly and has it within the first 50 files. There are no command-line flags: wayback-archive --help prints Error: WAYBACK_URL environment variable is required. Settings come from the environment or a .env file in the working directory. See How it works for the steps in detail and Configuration for every variable.
What it does not do¶
- It reads the Wayback Machine only. A live site is not a valid
WAYBACK_URL; for live sites usewget,httrackor web-mirror. - It never fetches the archived site's own domain, or any other host, from the live Internet, with three exceptions: Google Fonts,
code.jquery.comand Squarespace's CDN, and only for a file the archive lacks. Other external links stay external in the copy. - It cannot rebuild what the archive never captured. A page missing at every nearby timestamp is counted under
Files failedand its links stay broken. - It has no resume. A run interrupted with Ctrl-C leaves what it wrote in
OUTPUT_DIR; the next run starts from the first URL again. - It does not run JavaScript. Content a page loaded at runtime from an API is not discovered unless its URL appears in the page's HTML, CSS, JavaScript or
data-*attributes.
Privacy¶
- The tool contacts
web.archive.org, and only when a file is missing there,fonts.googleapis.comandfonts.gstatic.comfor a Google Font,code.jquery.comfor jQuery, or the Squarespace CDN for a file hosted there. It sends nothing anywhere else and has no telemetry. - Analytics and tracker scripts and ads are removed while the pages are rewritten (
REMOVE_TRACKERSandREMOVE_ADS, both on by default). External iframes, such as an embedded video, stay in the copy and load from the network unlessREMOVE_EXTERNAL_IFRAMES=true.REMOVE_CLICKABLE_CONTACTS=true(the default) also pointstel:andmailto:links at#. - The output is plain files on your disk. Nothing is uploaded.
Getting help¶
- If a page looks wrong, read Troubleshooting first: fonts, missing icons and libraries that do not load are the usual causes.
- Open an issue with the
WAYBACK_URL, the tail of the run's output including theDownload Complete!block, and your Python version. - Releases list what changed in each version.
- To send a fix, read Development.
License¶
Wayback-Archive is released under the GPL-3.0-or-later license.