In efforts to block AI companies' web crawlers, publishers are also not allowing Wayback Machine to take snapshots of their content, fearing that it could later be scraped from the pages of the archiving library.

The Wayback Machine is a project of the Internet Archive. It sends out web crawlers to take snapshots of the internet, creating a digital library of web pages. But now, some news publications are blocking its crawlers over concerns that AI companies will access the Wayback Machine’s publicly available archive and then train their AI models with the content.
Marketplace’s Stephanie Hughes talked about this with Andrew Deck at Harvard's Nieman Lab.
“News publishers limit Internet Archive access due to AI scraping concerns” from Nieman Lab
“The Internet's Most Powerful Archiving Tool Is in Peril” from Wired