How it works¶
- Initial download -- Fetches the main page from the Wayback Machine. If that capture redirects to another host (a site that moved domains), the tool warns and archives the host that was actually captured
- Link extraction -- Parses HTML to find all referenced assets (links, images, CSS, JS)
- CSS processing -- Extracts font URLs, background images, and
@importstatements; downloads Google Fonts locally; detects corrupted font files - JS processing -- Extracts dynamically loaded resources from JavaScript
- Data attributes -- Scans
data-*attributes for additional asset URLs - Iterative crawling -- Continues discovering and downloading resources until the queue is empty
- Timeframe fallback -- For 404 responses, tries up to three other timestamps: a day either side and a week earlier. Wayback already answers any timestamp with the nearest capture, so closer probes would land on the same answer. A capture that is an archived error (403, 5xx) or a Cloudflare challenge page is replaced by the nearest capture with status 200, found with one query to the Wayback CDX index; when the start page is replaced, the rest of the site is fetched around that capture's timestamp
- URL rewriting -- Converts all URLs to relative paths for offline serving
- Preservation -- Maintains icon groups, button links, and cookie consent functionality
Where requests go¶
Every file is requested from the Wayback Machine (web.archive.org). The tool never fetches the archived site's own domain, or any other host, from the live Internet: the domain may belong to someone else today, and what it serves now is not the archive.
The one exception is a short list of well-known CDN hosts. When the Wayback Machine does not have a file on one of these, the tool fetches it live, without following redirects:
- Google Fonts:
fonts.googleapis.com,fonts.gstatic.com - jQuery:
code.jquery.com(also the source of the replacement when a site's ownjquery.min.jsis missing: the version its URL names, or 3.7.1) - Squarespace:
static1.squarespace.com,static.squarespace.com,images.squarespace-cdn.com,sqspcdn.comand its subdomains
A host matches only exactly or as a subdomain (assets.sqspcdn.com), never by containing the name (sqspcdn.com.example.net, sqspcdn.com@10.0.0.1). Anything else the Wayback Machine does not have is reported as failed.
Project structure¶
Wayback-Archive/
wayback_archive/ # Main package
__init__.py
__main__.py
cli.py # CLI entry point
config.py # Environment variable configuration
downloader.py # Core download and processing engine
config/
requirements.txt # Runtime dependencies
requirements-dev.txt # Development dependencies
tests/ # Test suite
docs/ # Documentation
pyproject.toml # Package metadata, build and pytest configuration
LICENSE # GPL-3.0
README.md
Dependencies¶
| Package | Purpose |
|---|---|
| requests | HTTP client |
| beautifulsoup4 | HTML parsing |
| lxml | Fast HTML/XML parser |
| minify-html | HTML minification |
| cssmin | CSS minification |
| rjsmin | JS minification |
| Pillow | Image optimization |
| python-dotenv | .env file support |