Wayback Machine & Web Archiving – Internet Archive Help Center https://help.archive.org How can we help you? Fri, 26 Jun 2026 21:40:54 +0000 en-US hourly 1 https://wordpress.org/?v=6.8 https://help.archive.org/files/2024/03/cropped-Internet-Archive-Logo-White-on-Black-520x520-1-32x32.png Wayback Machine & Web Archiving – Internet Archive Help Center https://help.archive.org 32 32 FAQ: Publishers Blocking the Wayback Machine https://help.archive.org/help/faq-publishers-blocking-the-wayback-machine/ Fri, 15 May 2026 18:53:26 +0000 https://help.archive.org/?p=1968 Nieman Lab first reported that some publishers and news organizations have begun blocking the Internet Archive’s Wayback Machine from preserving and providing access to archived versions of their websites. 

Since then, journalists, digital rights advocates, historians, librarians, and researchers have raised alarms about the long-term consequences of limiting web preservation.

What is the Wayback Machine?

The Internet Archive is a nonprofit research library with a mission of providing Universal Access to All Knowledge. The Wayback Machine is a service of the Internet Archive that allows people to visit archived versions of web pages. The Wayback Machine has been archiving the web since 1996, helping preserve the historical record of the internet. Learn more about the Wayback Machine.

Why are publishers blocking the Wayback Machine?

As reported by journalists Andrew Deck and Hanaa’ Tameez in Nieman Lab, publishers say they are concerned about AI scraping and unauthorized reuse of their content by generative AI companies.

Some organizations have responded by broadly blocking web archiving systems like the Wayback Machine out of concern that archived material could be accessed or reused by AI systems.

Nicholas Thompson, CEO of The Atlantic, published a video explaining why The Atlantic has blocked the Wayback Machine from preserving its web content, saying:

“[The Wayback Machine] is an amazing resource…But, on the other hand, if you put everything up there, it will all not only go into the big AI companies, but to everybody who’s building some kind of subsidiary AI.”

But those concerns are unfounded, explains journalist Andrew Deck in an interview with Marketplace Tech:

“I think it’s important to say that in our conversations with news publishers, a lot of them were taking this action preemptively out of a fear of proxy scraping rather than direct evidence that it has happened to them already. None of the publishers were able to point to a particular AI company or other kinds of direct evidence that their content had already been scraped by the Wayback Machine.”

Michael Nelson, a computer scientist at Old Dominion University, has described the Wayback Machine and web archiving as “collateral damage” as content owners restrict access over AI scraping concerns. 

That sentiment was echoed by Mark Graham, director of the Wayback Machine, on the Future Knowledge podcast episode, “Preserving the Web in the Age of AI”:

“The Wayback Machine is collateral damage caught up in the conflict between AI companies and publishers.”

Who uses the Wayback Machine?

The Wayback Machine is widely used by journalists, fact-checkers, researchers, lawyers, courts, librarians, academics, and even publishers themselves.

More than 100 news articles every month reference, cite, or rely on material preserved by the Wayback Machine to verify claims, recover deleted information, or provide historical context.

In a Future Knowledge podcast interview, Mark Graham recalled a conversation at The New York Times:

“A senior researcher came up to me and said, ‘Oh my God, Mark, thank you so much for the Wayback Machine. We use you all the time. There is material available that we’ve used from the Wayback Machine that we can’t even find in our own archives.’”

Journalist Rachel Maddow also publicly defended the archive:

“The Internet Archive is a national treasure. I use it daily, and have for many, many years. I cannot imagine doing the work I do without it.”

What are critics saying about the blocking?

A broad coalition of journalists, digital rights advocates, and internet historians has warned that blocking the Wayback Machine—and web archiving tools like it—could have serious long-term consequences.

In Techdirt, Mike Masnick argued that publishers may regret these blocking decisions because they undermine preservation of the public record.

Joe Mullin, senior policy analyst at the Electronic Frontier Foundation, similarly warned that blocking the Internet Archive “won’t stop AI” but could erase important historical records from the web.

Meanwhile, coverage by journalist Kate Knibbs in Wired brought broader public attention to the issue and the risks facing digital preservation infrastructure.

Have journalists spoken out in support of the Wayback Machine?

Yes.

More than 200 journalists signed a public statement applauding the Internet Archive’s role in preserving the public record.

In response, Mark Graham published a public thank-you letter:

“Your support for the Wayback Machine sends a clear message: preserving the record matters.”

Does the Internet Archive respect publisher concerns?

Yes.

The Internet Archive works with publishers and rights holders to balance preservation, access, and responsible stewardship of digital materials.

What about AI scraping?

As Mark Graham explained in Techdirt:

“The Wayback Machine is built for human readers. We use rate limiting, filtering, and monitoring to prevent abusive access, and we watch for and actively respond to new scraping patterns as they emerge.”

What’s at stake when publishers block the Wayback Machine?

Every day that preservation systems are blocked leaves holes in the public record of the web. If preservation systems are weakened or blocked at scale, future generations will lose access to major parts of our digital history.

The size of the problem is significant. A 2024 study by Pew Research Center found that 38% of webpages from 2013 were no longer accessible a decade later, with roughly a quarter of pages sampled across the decade disappearing entirely. But loss on the web is not inevitable. New analysis by Internet Archive data scientist Sawood Alam found that the Wayback Machine has preserved roughly 15% of those otherwise vanished pages, saving reporting, citations, and pieces of the historical record that would no longer exist online.

As Mike Masnick wrote in Techdirt: “Blocking the Internet Archive isn’t going to stop AI training. What it will do is ensure that significant chunks of our journalistic record and historical cultural context simply… disappear.”

And as Mark Graham wrote:

“Preserving the public record is not optional. It is essential infrastructure for a functioning democracy.”


Chronological Reading List of Key Articles


FIRST PUBLISHED: May 15, 2026 CDF
LAST UPDATED: June 26, 2026 CDF

]]>
Save Pages in the Wayback Machine https://help.archive.org/help/save-pages-in-the-wayback-machine/ Fri, 15 Mar 2024 20:55:12 +0000 https://help.archive.org/?p=1447 Many people have shown interest in making sure the Wayback Machine has copies of the web pages they care about most. These saved pages can be cited, shared, linked to – and they will continue to exist even after the original page changes or is removed from the web.

There are several ways to save pages and whole sites so that they appear in the Wayback Machine.  Here are 5 of them.

1. Save Page Now

Put a URL into the form, press the button, and we save the page.  You will instantly have a permanent URL for your page. Please note, this method only saves a single page, not the whole site.

At the moment, there are a few exceptions for this method – some sites prohibit crawling, a few have SSL (security) settings that make it break – but this method will work for most pages.  The feature saves the page you enter including the images and CSS.  It does not save any of the outlinks, and can’t be used to initiate a crawl of an entire web site. We do not keep your IP address, so your submission is anonymous.

2. Browser extensions and add-ons

Install the Wayback Machine Chrome extension in your browser.  Go to a page you want to archive, click the icon in your toolbar, and select Save Page Now. We will save the page and give you a permanent URL.

The same provisos from “Save Page Now” apply – there are some pages where it won’t work, and it only saves one page at a time.  One plus to installing the extension though is that now as you surf around, when you run into a missing page we will alert you if we have a saved copy.

More extensions, apps, and add-ons:

3. Wikipedia JavaScript Bookmarklet

Nobody loves a primary source more than a Wikipedia editor.  To that end, they offer a Wayback Machine JavaScript Bookmarklet that allows you to quickly save a web page from any browser.

4. Volunteer for Archive Team

Archive Team is an entirely volunteer driven group who are interested in saving Internet history.  Many of the sites and pages they save end up in the Wayback Machine.  Visit the Archive Team site to learn more about how to volunteer with them.

5. Sign up for an Archive-It Account

Archive-It is a subscription service provided by Internet Archive that allows you to run your own crawling projects without any technical expertise.  Tell us what to crawl and how often to crawl it, and we execute the crawl and put the results in the Wayback Machine.

screenshot of webpage with information in Archive-It as leading service for preserving and accessing digital cultural heritage

Archive-It is a paid subscription service with technical and web archivist support. This option is most appropriate for organizations that have a mandate to save certain types or categories of web content on a regular basis. If your institution is a current Archive-It partner, contact them for how you can contribute.

The Internet Archive has been saving web pages for 20 years.  This archive has been built by thousands of people, and we would like you to help.  Use one of the methods above to make sure we have the pages you care about.

]]>
Archive whole web sites https://help.archive.org/help/archive-whole-web-sites/ Fri, 15 Mar 2024 20:54:24 +0000 https://help.archive.org/?p=1445 Organizations interested in archiving entire web sites or creating large collections of content may want to explore our Archive-It service.

Archive-It is a subscription web archiving service from the Internet Archive that helps organizations to harvest, build, and preserve collections of digital content. Through our user friendly web application Archive-It partners can collect, catalog, and manage their collections of archived content with 24/7 access and full text search available for their use as well as their patrons. 

Individuals who wish to archive web pages may want to refer to this article: Save Pages in the Wayback Machine

Developers may wish to consult the Wayback Machine API documentation.

]]>
Can I rebuild my website using the Wayback Machine? https://help.archive.org/help/can-i-rebuild-my-website-using-the-wayback-machine/ Fri, 15 Mar 2024 20:53:39 +0000 https://help.archive.org/?p=1443 There are several 3rd party services that will help you rebuild a website from archives available via our Wayback Machine. Here are some that we are aware of:

We do not have direct experience working with any of these sites, so please investigate the services to see whether they will meet your individual needs.

]]>
Wayback Machine General Information https://help.archive.org/help/wayback-machine-general-information/ Fri, 15 Mar 2024 20:53:04 +0000 https://help.archive.org/?p=1441 What is the Wayback Machine?

The Internet Archive Wayback Machine is a service that allows people to visit archived versions of Web sites. Visitors to the Wayback Machine can type in a URL, select a date range, and then begin surfing on an archived version of the Web. Imagine surfing circa 1999 and looking at all the Y2K hype, or revisiting an older version of your favorite Web site. The Internet Archive Wayback Machine can make all of this possible.

What are the sources of your captures?

When you roll over individual web captures (that pop-up when you roll over the dots on the calendar page for a URL,) you may notice some text links shows up above the calendar, along with the word “why”. Those links will take you to the Collection of web captures associated with the specific web crawl the capture came from. Every day hundreds of web crawls contribute to the web captures available via the Wayback Machine. Behind each, there is a story about factors like who, why, when and how.

Why is the Internet Archive collecting sites from the Internet? What makes the information useful?

Most societies place importance on preserving artifacts of their culture and heritage. Without such artifacts, civilization has no memory and no mechanism to learn from its successes and failures. Our culture now produces more and more artifacts in digital form. The Archive’s mission is to help preserve those artifacts and create an Internet library for researchers, historians, and scholars. The Archive collaborates with institutions including the Library of Congress and the Smithsonian.

Where does the name come from?

The Wayback Machine is named in reference to the famous Mr. Peabody’s WABAC (pronounced way-back) machine from the Rocky and Bullwinkle cartoon show.

Who was involved in the creation of the Internet Archive Wayback Machine?

“The original idea for the Internet Archive Wayback Machine began in 1996, when the Internet Archive first began archiving the web. Now, five years later, with over 100 terabytes and a dozen web crawls completed, the Internet Archive has made the Internet Archive Wayback Machine available to the public. The Internet Archive has relied on donations of web crawls, technology, and expertise from Alexa Internet and others. The Internet Archive Wayback Machine is owned and operated by the Internet Archive.”

How was the Wayback Machine made?

Alexa Internet, in cooperation with the Internet Archive, has designed a three dimensional index that allows browsing of web documents over multiple time periods, and turned this unique feature into the Wayback Machine.

How do you archive dynamic pages?

There are many different kinds of dynamic pages, some of which are easily stored in an archive and some of which fall apart completely. When a dynamic page renders standard html, the archive works beautifully. When a dynamic page contains forms, JavaScript, or other elements that require interaction with the originating host, the archive will not contain the original site’s functionality.

Do you collect all the sites on the Web?

No, the Archive collects web pages that are publicly available. We do not archive pages that require a password to access, pages that are only accessible when a person types into and sends a form, or pages on secure servers. Pages may not be archived due to robots exclusions and some sites are excluded by direct site owner request.

Do you archive email? Chat?

No, we do not collect or archive chat systems or personal email messages that have not been posted to Usenet bulletin boards or publicly accessible online message boards.

Is there any personal information in these collections?

We collect Web pages that are publicly accessible. These may include pages with personal information.

Who has access to the collections? What about the public?

Anyone can access our collections through our website archive.org. The web archive can be searched using the Wayback Machine.

The Archive makes the collections available at no cost to researchers, historians, and scholars. At present, it takes someone with a certain level of technical knowledge to access collections in a way other than our website, but there is no requirement that a user be affiliated with any particular organization.

What is the Wayback Machine’s Copyright Policy?

The Internet Archive respects the intellectual property rights and other proprietary rights of others. The Internet Archive may, in appropriate circumstances and at its discretion, remove certain content or disable access to content that appears to infringe the copyright or other intellectual property rights of others. If you believe that your copyright has been violated by material available through the Internet Archive, please provide the Internet Archive Copyright Agent with the following information:

Identification of the copyrighted work that you claim has been infringed;
An exact description of where the material about which you complain is located within the Internet Archive collections;
Your address, telephone number, and email address;
A statement by you that you have a good-faith belief that the disputed use is not authorized by the copyright owner, its agent, or the law;
A statement by you, made under penalty of perjury, that the above information in your notice is accurate and that you are the owner of the copyright interest involved or are authorized to act on behalf of that owner;
Your electronic or physical signature.
The Internet Archive Copyright Agent can be reached as follows:

Internet Archive Copyright Agent
Internet Archive
300 Funston Ave.
San Francisco, CA 94118
Phone: 415-561-6767
Email: info at archive dot org

How can I help the Internet Archive and the Wayback Machine?

The Internet Archive actively seeks donations of digital materials for preservation. If you have digital materials that may be of interest to future generations, please let us know by sending an email to info at archive dot org. The Internet Archive is also seeking additional funding to continue this important mission. You can click the donate tab above or click here. Thank you for considering us in your charitable giving.

How do I contact the Internet Archive?

All questions about the Wayback Machine, or other Internet Archive projects, should be addressed to info@archive.org.

]]>
Using the Wayback Machine https://help.archive.org/help/using-the-wayback-machine/ Fri, 15 Mar 2024 20:51:36 +0000 https://help.archive.org/?p=1436 This introduction video provides an overview for how to use the Wayback Machine, including information about searching by URL or keyword, understanding provenance, and saving your own pages, along with other features.

Can I link to old pages on the Wayback Machine?

Yes! The Wayback Machine is built so that it can be used and referenced. If you find an archived page that you would like to reference on your Web page or in an article, you can copy the URL. You can even use fuzzy URL matching and date specification… but that’s a bit more advanced.

How can I use the Wayback Machine’s Site Search to find websites?

The Site Search feature of the Wayback Machine is based on an index built by evaluating terms from hundreds of billions of links to the homepages of more than 350 million sites. Search results are ranked by the number of captures in the Wayback and the number of relevant links to the site’s homepage.

Can I search the Archive?

Using the Internet Archive Wayback Machine, it is possible to search for the names of sites contained in the Archive (URLs) and to specify date ranges for your search. We hope to implement a full text search engine at some point in the future.

Why isn’t the site I’m looking for in the archive?

Some sites may not be included because the automated crawlers were unaware of their existence at the time of the crawl. It’s also possible that some sites were not archived because they were password protected, blocked by robots.txt, or otherwise inaccessible to our automated systems. Site owners might have also requested that their sites be excluded from the Wayback Machine.

How can I exclude or remove my site’s pages from the Wayback Machine?

If you would like to submit a request for archives of your site or account to be excluded from web.archive.org, send us a request to info@archive.org and indicate:

  • the URL or URLs of the material
  • the time period that you wish to have excluded
  • the time period during which you had control of the site or relevant user account (if applicable) and 
  • any other information that you think would be helpful for us to better understand your request. 

This will initiate a review by our team. We do not make any guarantees beforehand about the outcome of a request.

How can I use the Wayback Machine’s Site Search to find websites?

The Site Search feature of the Wayback Machine is based on an index built by evaluating terms from hundreds of billions of links to the homepages of more than 350 million sites. Search results are ranked by the number of captures in the Wayback and the number of relevant links to the site’s homepage.

How can I get a copy of the pages on my Web site? If my site got hacked or damaged, could I get a backup from the Archive?

Our terms of use do not cover backups for the general public. However, you may use the Internet Archive Wayback Machine to locate and access archived versions of a site to which you own the rights. We can’t guarantee that your site has been or will be archived. We can no longer offer the service to pack up sites that have been lost.

Can I add pages to the Wayback Machine?

On https://archive.org/web you can use the “Save Page Now” feature to save a specific page one time. This does not currently add the URL to any future crawls nor does it save more than that one page. It does not save multiple pages, directories or entire sites.

Where is the rest of the archived site? Why am I getting broken or gray images on a site?

Broken images occur when the images are not available on our servers. Usually this means that we did not archive them.

You can tell if the image or link you are looking for is in the Wayback Machine by entering the image or link’s URL into the Wayback Machine search box. Whatever archives we have are viewable in the Wayback Machine.

The best way to see all the files we have archived of the site is: http://web.archive.org/*/www.yoursite.com/*

There is a 3-10 hour lag time between the time a site is crawled and when it appears in the Wayback Machine.

Why are some sites harder to archive than others?

If you look at our collection of archived sites, you will find some broken pages, missing graphics, and some sites that aren’t archived at all. Some of the things that may cause this are:

Robots.txt — A site’s robots.txt document may have prevented the crawling of a site.
Javascript — Javascript elements are often hard to archive, but especially if they generate links without having the full name in the page. Plus, if javascript needs to contact the originating server in order to work, it will fail when archived.
Server side image maps — Like any functionality on the web, if it needs to contact the originating server in order to work, it will fail when archived.
Orphan pages — If there are no links to your pages, the robot won’t find it (the robots don’t enter queries in search boxes.)
As a general rule of thumb, simple html is the easiest to archive.
Can I find sites by searching for words that are in their pages?

No, at least not yet. Site Search for the Wayback Machine will help you find the homepages of sites, based on words people have used to describe those sites, as opposed to words that appear on pages from sites.

Can I still find sites in the Wayback Machine if I just know the URL?

Yes, just enter a domain or URL the way you have in the past and press the “Browse History” button.

Why are some of the dots on the calendar page different colors?

We color the dots, and links, associated with individual web captures, or multiple web captures, for a given day. Blue means the web server result code the crawler got for the related capture was a 2nn (good); Green means the crawlers got a status code 3nn (redirect); Orange means the crawler got a status code 4nn (client error), and Red means the crawler saw a 5nn (server error). Most of the time you will probably want to select the blue dots or links.

How does the Wayback Machine behave with Javascript turned off?

If you have Javascript turned off, images and links will be from the live web, not from our archive of old Web files.

How did I end up on the live version of a site? or I clicked on X date, but now I am on Y date, how is that possible?

Not every date for every site archived is 100% complete. When you are surfing an incomplete archived site the Wayback Machine will grab the closest available date to the one you are in for the links that are missing. In the event that we do not have the link archived at all, the Wayback Machine will look for the link on the live web and grab it if available. Pay attention to the date code embedded in the archived url. This is the list of numbers in the middle; it translates as yyyymmddhhmmss. For example in this url http://web.archive.org/web/20000229123340/http://www.yahoo.com/ the date the site was crawled was Feb 29, 2000 at 12:33 and 40 seconds.

You can see a listing of the dates of the specific URL by replacing the date code with an asterisk (*), ie: http://web.archive.org/*/www.yoursite.com

How do I cite Wayback Machine urls in MLA format?

This question is a newer one. We asked MLA to help us with how to cite an archived URL in correct format. They did say that there is no established format for resources like the Wayback Machine, but it’s best to err on the side of more information. You should cite the webpage as you would normally, and then give the Wayback Machine information. They provided the following example: McDonald, R. C. “Basic Canary Care.” _Robirda Online_. 12 Sept. 2004. 18 Dec. 2006 [http://www.robirda.com/cancare.html]. _Internet Archive_. [ http://web.archive.org/web/20041009202820/http://www.robirda.com/cancare.html]. They added that if the date that the information was updated is missing, one can use the closest date in the Wayback Machine. Then comes the date when the page is retrieved and the original URL. Neither URL should be underlined in the bibliography itself. Thanks MLA!

How can I get pages authenticated from the Wayback Machine? How can I use the pages in court? While the Wayback Machine tool was not expressly designed with legal use in mind, we receive regular requests for certified records for use in legal proceedings. Our affidavit request procedure can be found here. Please review that information including our standard affidavit and the legal request FAQ section linked there to prior to contacting us.

Some sites are not available because of robots.txt or other exclusions. What does that mean?

Such sites may have been excluded from the Wayback Machine due to a robots.txt file on the site or at a site owner’s direct request.

How can I get my site included in the Wayback Machine?

Much of our archived web data comes from our own crawls or from Alexa Internet’s crawls. Neither organization has a “crawl my site now!” submission process. Internet Archive’s crawls tend to find sites that are well linked from other sites. The best way to ensure that we find your web site is to make sure it is included in online directories and that similar/related sites link to you.

Alexa Internet uses its own methods to discover sites to crawl. It may be helpful to install the free Alexa toolbar and visit the site you want crawled to make sure they know about it.

Regardless of who is crawling the site, you should ensure that your site’s ‘robots.txt’ rules and in-page META robots directives do not tell crawlers to avoid your site.

What is the Archive-It service of the Internet Archive’s Wayback Machine?

For information on the Archive-It subscription service that allows institutions to build and preserve collections of born digital content, see https://www.archive.org/about/faqs.php#Archive-It.

]]>
Why can’t I see the Web page I archived recently? https://help.archive.org/help/why-cant-i-see-the-web-page-i-archived-yesterday/ Wed, 02 Mar 2022 20:52:27 +0000 https://help.blog.archive.org/?p=370 The Wayback Machine can sometimes experience delays in registering snapshots made using the Save Page Now tool. Our website replays pages only after they’ve been saved (indexed) by our system. We use several different indexes and each one covers a different time period.

Sometimes a web page that gets added quickly to our near-real-time index used by our Save Page Now service may be removed from that index before it gets stored in one of our longer-term indexes. When that happens, you might see a page available to view for a short time—maybe a few hours or days—but then it disappears for a while. It will become available again once it’s added to a longer-term index.

Rest assured that the snapshot was still captured at the time you requested— we appreciate your patience while you wait! If you are still unable to view a snapshot you took after some time, please reach out to info@archive.org with relevant details.

]]>