Skip to content
RU
← All articles

Wayback Machine Guide: Find, Save and Remove Archived Pages

A web archive calendar of captures and an old version of a site in another tab

The Wayback Machine is the Internet Archive's public web archive at web.archive.org: it stores dated snapshots of web pages so you can see how any URL looked in the past. Type a domain or page address into its search box, pick a year, then click a highlighted day in the calendar to open that capture.

What the Wayback Machine actually stores

The archive does not keep websites as a whole. It keeps captures: one URL, one moment in time, plus whatever images, stylesheets and scripts the crawler managed to fetch alongside it. When a resource is missing from a capture, the replay engine borrows the closest copy it has from another date, or leaves a hole. That is why an old homepage can show this year's logo on a layout from five years earlier. Treat every snapshot as a reconstruction, not a photograph.

This also explains the key difference from a search engine cache. A cache held one copy — the last crawl — and disappeared along with the live page; Google has retired its public cache links altogether. The Wayback Machine keeps the history, and that history survives after a domain expires or changes hands.

Typical reasons people open it:

  • recovering text or images that were deleted from your own site;
  • rebuilding the old URL structure before a migration so every legacy address gets a redirect;
  • checking what a domain was used for before you buy it;
  • documenting what a company published on a given date, as a supporting source.

How to browse an old version of a website

  1. Go to web.archive.org and paste a domain (example.com) or a full page URL into the search field.
  2. A histogram of years appears above the calendar; taller bars mean more captures. Click a year.
  3. Days with captures are marked with circles. The colour reflects the HTTP status the crawler received: blue for 2xx, green for 3xx redirects, orange for 4xx client errors, red for 5xx server errors.
  4. Hover over a day and choose a timestamp. The capture opens with the Wayback toolbar on top, which lets you step to neighbouring snapshots.

Links inside a capture are rewritten to point back into the archive, so you can click around the old site and the replay engine will serve the nearest capture of each page. When nothing exists, you will see "Hrm. Wayback Machine has not archived that URL."

The results page also has tabs worth knowing: Changes highlights differences between two captures, Summary breaks down archived file types by year, and URLs lists every address captured under the domain.

Jumping straight to a date: the URL patterns

Everything in the interface maps to a predictable URL scheme. Timestamps are 14 digits, YYYYMMDDhhmmss in UTC, and can be truncated.

URLResult
web.archive.org/web/example.commost recent capture
web.archive.org/web/2016/example.comcapture closest to that year
web.archive.org/web/20160315120000/example.com/pricingcapture closest to 15 March 2016, 12:00 UTC
web.archive.org/web/2016*/example.comcalendar view for 2016
web.archive.org/web/*/example.com/*every captured URL on the domain
web.archive.org/web/20160315id_/example.comoriginal HTML with no toolbar and no rewritten links

The trailing-wildcard pattern is the one most people miss. It turns the archive into an inventory of every page the domain ever exposed, with first and last capture dates — forgotten landing pages, old PDFs and deleted blog posts included.

The id_ modifier matters when you recover content. Without it, saved HTML is full of archive scripts and rewritten URLs; with it, you get the markup as the original server sent it.

Querying the archive from a terminal

Two public APIs cover most needs. The availability endpoint tells you whether a capture exists near a date:

curl -s "https://archive.org/wayback/available?url=example.com&timestamp=20160101"

The JSON response contains archived_snapshots.closest with the snapshot URL, timestamp and status; an empty archived_snapshots object means there is nothing. The CDX API returns the full capture index. This lists unique URLs that ever answered 200:

curl -s "https://web.archive.org/cdx/search/cdx?url=example.com/*&fl=original&filter=statuscode:200&collapse=urlkey" > legacy-urls.txt

Add from=2015&to=2018 to limit the period or output=json for structured output. On Windows, Invoke-RestMethod in PowerShell handles the same requests and parses the JSON for you.

Why a page is missing from the archive

  • It was never crawled. Coverage is selective. Homepages of well-linked sites are captured often; deep pages on small sites may never be.
  • It is rendered by JavaScript. Single-page apps that fetch content after load often replay as an empty shell or an endless spinner.
  • It sits behind a login. Dashboards, private groups and most social feeds replay as a sign-in form.
  • The owner requested exclusion. You will see "Sorry. This URL has been excluded from the Wayback Machine."
  • The server refused the crawler. A 403, a 503 or bot filtering on the capture date shows up as an orange or red mark instead of content.
  • The URL differs. Query strings count as part of the address. Use the /* wildcard and filter the list.

Video is a special case: YouTube pages are often archived, but the stream itself usually is not, so the player in a capture will not play.

Saving a page yourself with Save Page Now

You do not have to wait for the crawler. The Save Page Now form on web.archive.org archives a URL on demand, and the same thing works as a link:

https://web.archive.org/save/https://example.com/terms

Signed-in users get extra options such as capturing outlinks, saving error pages and taking a screenshot. The official browser extension adds a one-click save button. A sensible habit for site owners: capture your terms, pricing and privacy pages before every significant change, so you can later prove what they said and when.

How to get your site removed

The Internet Archive's help centre directs removal requests to info@archive.org. Include the exact URLs or the whole domain, the time range, proof that you control the site (for example, writing from an address on that domain) and the reason. A robots.txt rule aimed at the archive's crawler used to be the common approach, but the Archive's handling of robots.txt has changed over the years, and a directive alone does not guarantee that existing captures are hidden. A direct request is the reliable route.

Domain due diligence and SEO

Before buying an expired or aftermarket domain, read its archive history. Red flags include gambling or pharma content unrelated to your niche, doorway pages generated by the thousand, a sudden switch of language or country, and long stretches of parking pages that indicate the name already dropped once. The archive shows content only; registration dates and ownership come from WHOIS. See how to find out who owns a domain and the guide to expiring domains.

For SEO, the CDX list of legacy URLs is the backbone of a migration redirect map: every old address should return a 301 to its closest new equivalent, not a 404. The Changes tab is equally useful after a traffic drop — comparing captures from before and after shows exactly which titles, copy blocks or internal links disappeared. The full sequence is in our website migration checklist, and redirect choice is covered in 301 vs 302 redirects.

Wayback Machine alternatives

OptionHow copies are madeBest forLimits
Wayback Machineown crawlers plus on-demand saveslong history, calendar, CDX API, diffingpatchy on JavaScript-heavy pages; owners can request exclusion
archive.today (archive.ph)on demand onlyfaithful rendering of dynamic pages, screenshot includedno automatic crawling, so no history unless someone saved it
Search engine cachelast crawlquick look at a recent version, where still offeredsingle copy, no history; Google no longer provides it
Your own backupshosting snapshots or version controlcomplete site including database and server logiconly exists if you set it up in advance

How to check a domain alongside its history

  • WHOIS lookup — creation date, registrar, expiry and status codes. A recent creation date on a domain with a long archive history means it was dropped and re-registered.
  • Redirect checker — confirms that legacy URLs from your CDX export now return a single 301 hop.
  • HTTP header check — what the live site returns today, for comparison with the status the crawler recorded.

Official documentation: Using the Wayback Machine and Wayback Machine General Information.

Frequently asked questions

Is the service free?

Yes. Browsing captures, Save Page Now and the public APIs cost nothing; an account only unlocks extra save options.

How often are sites archived?

There is no fixed schedule. Popular pages are captured frequently, obscure ones rarely or never. If a specific date matters, save the page yourself.

Can I download an entire site from the archive?

You can reconstruct its static pages by combining the CDX URL list with id_ requests, but forms, search and anything server-side cannot be recovered. Keep request rates modest.

Can a snapshot be used as evidence?

It is widely cited as supporting evidence, but admissibility depends on the jurisdiction and the court. Where it matters, ask a lawyer how captures should be authenticated.

Why does an archived page look broken?

Missing CSS, images or scripts were either never captured or came from a different date. Try another timestamp, or use archive.today for pages that depend on JavaScript.

Check your website right now

Check your domain →
More articles: Tools
Tools
Batch URL Checking: Automating Website Monitoring
11.03.2026 · 542 views
Tools
Website Builders 2026: What You Give Up and How to Migrate
21.07.2026 · 213 views
Tools
What Is an MCP Server and Why It Matters
15.06.2026 · 194 views
Tools
How to Give Claude and Cursor Web Diagnostic Tools
15.06.2026 · 175 views