Features¶
Core¶
- Full website download -- HTML, CSS, JS, images, fonts, and all linked assets
- Recursive link discovery -- Automatically follows links in HTML, CSS, and JS files
- Smart URL rewriting -- Converts all links to relative paths for local serving
- Timeframe fallback -- Tries up to three other Wayback Machine timestamps (a day either side and a week earlier) when a resource returns 404
- Bad capture fallback -- When a capture is an archived error (403, 5xx) or a Cloudflare challenge page, uses the nearest capture with status 200 instead, found with one query to the Wayback CDX index
- Real-time progress logging -- Displays download status and file processing as it happens
Asset Handling¶
- Google Fonts support -- Downloads Google Fonts CSS and font files locally, fixing CORS issues
- Font corruption detection -- Identifies and removes corrupted font files (HTML error pages served as fonts)
- CDN fallback -- When the Wayback Machine lacks a file hosted on Google Fonts,
code.jquery.comor the Squarespace CDN, fetches it from that CDN; a missingjquery.min.jsis replaced fromcode.jquery.comwith the version its URL names (jquery-1.7.2/,?ver=3.6.0;1.9is also tried as1.9.0), or 3.7.1 when it names none orcode.jquery.comdoes not have that version (a WordPress?ver=is often WordPress's own version). Nothing else is fetched live (see How it works) - Data attribute processing -- Processes
data-*attributes containing URLs (videos, images, etc.)
Preservation¶
- Icon group preservation -- Preserves all links in icon groups (social media, contact icons)
- Button link preservation -- Maintains styling and functionality of button links
- Cookie consent preservation -- Keeps cookie consent popups and functionality intact
Optimization¶
- HTML minification -- Uses
minify-html(Python 3.14+ compatible) - JS/CSS minification -- Optional JavaScript and CSS minification via
rjsminandcssmin - Image compression -- Optional image optimization with Pillow
- Tracker/ad removal -- Strips analytics and ads by default, and external iframes when asked
- Link cleanup -- Configurable external link removal with anchor preservation options
- www/non-www normalization -- Drops
www.from the archived site's host by default, or adds it withMAKE_WWW
Why Wayback-Archive?¶
| Capability | Wayback-Archive | wget | httrack |
|---|---|---|---|
| Wayback Machine URL rewriting | Yes | No | No |
| Wayback artifact cleanup | Yes | No | No |
| Timeframe fallback for 404s | Yes | No | No |
| Google Fonts localization | Yes | No | No |
| Font corruption detection | Yes | No | No |
| CDN fallback | Yes | No | No |
| HTML/CSS/JS minification | Yes | No | No |
| Tracker and ad removal | Yes | No | No |
data-* attribute processing |
Yes | No | No |
General-purpose tools like wget --mirror or httrack can download live websites, but they do not understand Wayback Machine URL structures, cannot clean up archive artifacts, and lack the specialized asset recovery that Wayback-Archive provides.