Archive.org General Information – Internet Archive Help Center https://help.archive.org How can we help you? Fri, 26 Jun 2026 21:40:54 +0000 en-US hourly 1 https://wordpress.org/?v=6.8 https://help.archive.org/files/2024/03/cropped-Internet-Archive-Logo-White-on-Black-520x520-1-32x32.png Archive.org General Information – Internet Archive Help Center https://help.archive.org 32 32 FAQ: Publishers Blocking the Wayback Machine https://help.archive.org/help/faq-publishers-blocking-the-wayback-machine/ Fri, 15 May 2026 18:53:26 +0000 https://help.archive.org/?p=1968 Nieman Lab first reported that some publishers and news organizations have begun blocking the Internet Archive’s Wayback Machine from preserving and providing access to archived versions of their websites. 

Since then, journalists, digital rights advocates, historians, librarians, and researchers have raised alarms about the long-term consequences of limiting web preservation.

What is the Wayback Machine?

The Internet Archive is a nonprofit research library with a mission of providing Universal Access to All Knowledge. The Wayback Machine is a service of the Internet Archive that allows people to visit archived versions of web pages. The Wayback Machine has been archiving the web since 1996, helping preserve the historical record of the internet. Learn more about the Wayback Machine.

Why are publishers blocking the Wayback Machine?

As reported by journalists Andrew Deck and Hanaa’ Tameez in Nieman Lab, publishers say they are concerned about AI scraping and unauthorized reuse of their content by generative AI companies.

Some organizations have responded by broadly blocking web archiving systems like the Wayback Machine out of concern that archived material could be accessed or reused by AI systems.

Nicholas Thompson, CEO of The Atlantic, published a video explaining why The Atlantic has blocked the Wayback Machine from preserving its web content, saying:

“[The Wayback Machine] is an amazing resource…But, on the other hand, if you put everything up there, it will all not only go into the big AI companies, but to everybody who’s building some kind of subsidiary AI.”

But those concerns are unfounded, explains journalist Andrew Deck in an interview with Marketplace Tech:

“I think it’s important to say that in our conversations with news publishers, a lot of them were taking this action preemptively out of a fear of proxy scraping rather than direct evidence that it has happened to them already. None of the publishers were able to point to a particular AI company or other kinds of direct evidence that their content had already been scraped by the Wayback Machine.”

Michael Nelson, a computer scientist at Old Dominion University, has described the Wayback Machine and web archiving as “collateral damage” as content owners restrict access over AI scraping concerns. 

That sentiment was echoed by Mark Graham, director of the Wayback Machine, on the Future Knowledge podcast episode, “Preserving the Web in the Age of AI”:

“The Wayback Machine is collateral damage caught up in the conflict between AI companies and publishers.”

Who uses the Wayback Machine?

The Wayback Machine is widely used by journalists, fact-checkers, researchers, lawyers, courts, librarians, academics, and even publishers themselves.

More than 100 news articles every month reference, cite, or rely on material preserved by the Wayback Machine to verify claims, recover deleted information, or provide historical context.

In a Future Knowledge podcast interview, Mark Graham recalled a conversation at The New York Times:

“A senior researcher came up to me and said, ‘Oh my God, Mark, thank you so much for the Wayback Machine. We use you all the time. There is material available that we’ve used from the Wayback Machine that we can’t even find in our own archives.’”

Journalist Rachel Maddow also publicly defended the archive:

“The Internet Archive is a national treasure. I use it daily, and have for many, many years. I cannot imagine doing the work I do without it.”

What are critics saying about the blocking?

A broad coalition of journalists, digital rights advocates, and internet historians has warned that blocking the Wayback Machine—and web archiving tools like it—could have serious long-term consequences.

In Techdirt, Mike Masnick argued that publishers may regret these blocking decisions because they undermine preservation of the public record.

Joe Mullin, senior policy analyst at the Electronic Frontier Foundation, similarly warned that blocking the Internet Archive “won’t stop AI” but could erase important historical records from the web.

Meanwhile, coverage by journalist Kate Knibbs in Wired brought broader public attention to the issue and the risks facing digital preservation infrastructure.

Have journalists spoken out in support of the Wayback Machine?

Yes.

More than 200 journalists signed a public statement applauding the Internet Archive’s role in preserving the public record.

In response, Mark Graham published a public thank-you letter:

“Your support for the Wayback Machine sends a clear message: preserving the record matters.”

Does the Internet Archive respect publisher concerns?

Yes.

The Internet Archive works with publishers and rights holders to balance preservation, access, and responsible stewardship of digital materials.

What about AI scraping?

As Mark Graham explained in Techdirt:

“The Wayback Machine is built for human readers. We use rate limiting, filtering, and monitoring to prevent abusive access, and we watch for and actively respond to new scraping patterns as they emerge.”

What’s at stake when publishers block the Wayback Machine?

Every day that preservation systems are blocked leaves holes in the public record of the web. If preservation systems are weakened or blocked at scale, future generations will lose access to major parts of our digital history.

The size of the problem is significant. A 2024 study by Pew Research Center found that 38% of webpages from 2013 were no longer accessible a decade later, with roughly a quarter of pages sampled across the decade disappearing entirely. But loss on the web is not inevitable. New analysis by Internet Archive data scientist Sawood Alam found that the Wayback Machine has preserved roughly 15% of those otherwise vanished pages, saving reporting, citations, and pieces of the historical record that would no longer exist online.

As Mike Masnick wrote in Techdirt: “Blocking the Internet Archive isn’t going to stop AI training. What it will do is ensure that significant chunks of our journalistic record and historical cultural context simply… disappear.”

And as Mark Graham wrote:

“Preserving the public record is not optional. It is essential infrastructure for a functioning democracy.”


Chronological Reading List of Key Articles


FIRST PUBLISHED: May 15, 2026 CDF
LAST UPDATED: June 26, 2026 CDF

]]>
SFLan Information https://help.archive.org/help/sflan-information/ Fri, 15 Mar 2024 20:22:41 +0000 https://help.archive.org/?p=1380 How can I connect to SFLan?
With a laptop: Be in the vicinity of a SFLan node. Associate with it: The SSID is sflanNN, where NN is the number of node, e.g. sflan13. No WEP. You’ll get an IP number assigned via DHCP. With a house: Contact us at info at archive dot org. (Please include your address and a phone number.) Find out if you have line of sight to another SFLan node, buy a node, and we’ll put it on your roof.

What about IP addresses?
SFLan uses real, routable IP addresses. These are usually given out dynically via DHCP. The nodes themselves use static addresses. We can also assign static addresses for servers. For the techies: We use tunneling, layer 2 and layer 3 bridging in parts on the network to make it all appear as a “flat” LAN. There are pros and cons about this approach. It has worked best for us so far. However, it is a moving target, and might change in the future.

I still have more questions, what should I do?
SFLan is a work in progress. If you have more questions, try the SFLan forum. If you still need help, write to info at archive dot org.

I live at 123 Main St at Crossing; do I have line of sight access to a node?
You can try netstumbler or kismet to look for a SFLan ssid.

What is the cost of a node?
The nodes cost $1100, which includes the price of parts and installation. Discounts are potentially available depending on the location.

How can I get a node?
Send an email with your name, exact address and phone number to info at archive dot org. Be sure to write “SFLan node” (or something similar) in the subject line. The information will be passed on to our fantastic installation team who will contact you.

If I get a node, can my neighbors connect also?
Yes, a SFLan node can connect your neighbors and co-condo association members.

What is included in the node?
Most of our nodes are composed of two radios, but some have three. The components are in a weather tight box with a four foot coax cable and two antennas attached. The whole unit is mounted on your roof (generally) on a pole. There is a picture of our lovely 5’3″ spokesmodel holding one here: http://www.archive.org/iathreads/uploaded-files/AstridB-PICT0017.JPG

What are the power requirements of a node?
A node takes on average 5 watts.

What are the connection characteristics of the network?
There are no average characteristics, but 2MBs shared among 20 or so people would be an example.

What is the percentage of uptime?
SFLan is an experimental network, so the uptime varies. Right now uptime averages around 90% or more.

]]>
Archive.org Information https://help.archive.org/help/archive-org-information/ Fri, 15 Mar 2024 19:17:04 +0000 https://help.archive.org/?p=1324 What’s the significance of the Archive’s collections?
Societies have always placed importance on preserving their culture and heritage. But much early 20th-century media — television and radio, for example — was not saved.  The Library of Alexandria  — an ancient center of learning containing a copy of every book in the world — disappeared when it was burned to the ground.

What are your fees?
At this time we have no fees for uploading and preserving materials. We estimate that permanent storage costs us approximately $2.00US per gigabyte. While there are no fees we always appreciate donations to offset these costs.

Do you backup my files?
Yes. We duplicate/backup all files at various locations

How long will you store files?
As an archive our intention is to store and make materials in perpetuity.

What languages are supported by Archive.org?
Archive.org supports all metadata about items in just about any language so long as the characters are UTF8 encoded.

What is a “view”?
A “view” used to be called a “download” on archive.org. How are “views” counted?
archive.org calculates a view as: one action (read a book, download a file, watch a movie, etc.), per day, per IP Address. So, for each item page, using multiple files or accessing from multiple accounts in a single day will only count as one view.

How often are views counted?
View are not counted for at least 14 days after they occur. Collection counts shown in the graph on the “About” page are updated monthly.

What is GDPR?
The EU General Data Protection Regulation (GDPR) builds upon and modernizes existing EU Data Protection and Privacy rules. We operate under a terms of service, copyright policy and privacy policy, that are available here. Please read these. 

As a library, the Internet Archive has, in the words of the GDPR, a “legitimate interest” in building collections, providing permanent public access, and maintaining archival integrity. 
In general, the Internet Archive, well,  archives  digital and physical materials that we collect ourselves or is contributed by users based on their uses of our services. This includes collecting available provenance information such as uploader and dates. We try to keep everything forever and try to make everything available publicly, and if not publicly, then at least to researchers, historians, and scholars. That is how we see our job.
Updating, deleting, and exporting account information is available to each user, not limited to residents of the EU. If you have uploaded things to the Internet Archive, you can find a list of them from the “my library” link on your settings page. From there, you can download the items and take them anywhere.

We use limited automated techniques to reduce spam and limit damage to our services. We hope to increase these in the future. Users that object may write to info@archive.org.
Our collections reside on servers in many locations in order to keep them safe, and data may be processed outside of the European Economic Area.

]]>
Internet Archive Statistics https://help.archive.org/help/internet-archive-statistics/ Fri, 15 Mar 2024 19:16:15 +0000 https://help.archive.org/?p=1322 What user stats do you keep and share?

The only users stats we track are the “views” of items on the site.

Where are the stats?

For collections they are viewable in a chart form in the “About” tab on a collection page. These numbers represent views in all the items in that collection. These are updated daily. For items they are shown on the right side of the details page. These are updated daily. Search results pages also show the “views” to the left of the page title. These numbers may differ from those on item and collection pages because they are updated monthly rather than daily.

What is a “view”?

A “view” used to be called a “download” on archive.org. How are “views” counted?
archive.org calculates a view as: one action (read a book, download a file, watch a movie, etc.), per day, per IP Address. So, for each item page, using multiple files or accessing from multiple accounts in a single day will only count as one view.

How often are views counted?

Item pages are updated daily so the current number would reflect the count through the previous day. Collection counts shown in the graph on the “About” page are updated monthly.

]]>
Archive.org site architecture and glossary https://help.archive.org/help/archive-org-site-architecture-and-glossary/ Fri, 15 Mar 2024 19:14:29 +0000 https://help.archive.org/?p=1320 Archive.org contains millions of items of a wide variety of mediatypes such as texts, audio, movies, software, images and more. The site technically is flat. The metadata gives it the appearance of a hierarchical structure. Here are some of the main features of the site.

Site Organization

Mediatype
The site is organized into silos with the top of each silo being a mediatype. There is texts, audio, movies, software, images, data and web. Each child collection inherits that mediatype from its parent so, for example, all collections under movies will by default add new items as mediatype=movies. Several mediatypes have a player unique to that mediatype. Texts has a bookreader, Audio and Movies a player, Software has emulation items, Image has a slideshow and web has the Wayback Machine.

Collection
A collection is a group of item pages organized under a collections page. There can be collections within collections. Many are created for scanning partners or by the Internet Archive but uploaders may also request collections be made for their items once they have created at least 50. Each item in a collection automatically inherits the collections parent as well so items appear not only in their collection but in the one above it.

Item
An item is a page on the site with data and metadata. Items can be based on a single uploaded source file, like a book, or many source files, like a live concert with many songs. Items are created when a file is uploaded. 

Tasks
The archive system then runs a series of tasks and, depending on the mediatype and file format that was uploaded, creates derivative files. Some of these files are intended to be web-friendly so they will play in online players, some are metadata, some are so that the item can function of the site. You can see the log of these tasks by modifying an item’s /details/ URL to be /history/ instead. They are color coded. A red task means that something failed and may need admin attention.

Account pages
Accounts automatically have several pages associated with them; Favorites, Settings, Loans, Library. These allow you to see your activity on the site as well as create lists. These pages cannot be removed. The “@” user name cannot be modified once created.

Other terms used on the site:
identifier – The id of the item. It is the tail of the item’s URL
file format – the type of file. for example, mp4, zip, xml, flac. There are many, many file formats.
player – The streaming player that lets you experience audio, movies and some kinds of software
bookreader – The “player” for text items. You can see it by clicking the “fullscreen” icon in the upper right side of a text item’s page.
restricted – Some items are restricted from public use for a variety of reasons.
derive – The task that creates other files from the uploaded file
view – A view is the equivalent of a use whether it is a download or a play on the site. Archive.org counts views as one use, per item, per day, per IP address.
facets – these are the metadata options list on the left side of search results and collection pages. They allow you to narrow your results.

For more detailed definitions and explanations see the Technical Information page.

Do you backup my files?

Yes. We duplicate/backup all files at various locations.

]]>
Developer Resources https://help.archive.org/help/developer-resources/ Fri, 15 Mar 2024 18:57:07 +0000 https://help.archive.org/?p=1285 The Internet Archive Developers Portal is the best place to get up to date information, but we have included a few links here as shortcuts if you already know what you’re looking for:

]]>
Introductory Tour of Archive.org and its Collections https://help.archive.org/help/introductory-tour-of-archive-org-and-its-collections/ Fri, 15 Mar 2024 18:52:33 +0000 https://help.archive.org/?p=1275 The Internet Archive is a 501(c)(3) non-profit library founded to gather knowledge in digital formats and make it available to everyone in the world for free. But 80+ million objects (books, web pages, TV, movies, radio, concerts, etc.) can be a daunting library to navigate.

This video provides:

  • An overview of the major collections on archive.org and how to use them
  • Tips for how to efficiently find what you’re looking for
  • Pointers for where to find more information on your own

Make sure you’re ready to take notes, there’s a lot of information to share!

]]>
Rights https://help.archive.org/help/rights/ Fri, 15 Mar 2024 18:51:47 +0000 https://help.archive.org/?p=1273 What happens if I post content that infringes someone’s copyright?

If the Internet Archive is made aware of content that infringes someone’s copyright, we will remove it per our Copyright Policy.

We have a policy of terminating the accounts of users who we determine, in our discretion, to be “repeat infringers” of copyright.

If you believe material uploaded to archive.org was removed due to a copyright complaint as a result of mistake or misidentification, you may submit a counter-notice contesting the copyright complaint. Upon our receipt of a valid counter-notice, we may wait 10 to 14 days to restore the material, unless the copyright owner notifies us that it has initiated legal action against you.

To be effective, a counter-notice must contain the following information:

  1. The URL of the material that was removed (This was in the email informing you of the copyright complaint);
  2. A statement under penalty of perjury that you have a good faith belief that the material was removed as a result of mistake or misidentification. (“Under penalty of perjury I have a good faith belief that the material was removed or disabled as a result of mistake or misidentification of the material.”) 
  3. Your name, address, and telephone number, and a statement of consent to jurisdiction  (“I consent to the jurisdiction of Federal District Court for the judicial district in which my address is located, if in the United States, otherwise the Northern District of California where the Internet Archive is located, and I will accept service of process from the person who provided the copyright complaint or an agent of such person.”); and 
  4. Your physical or electronic signature.

Email or send your counter-notice to:

Internet Archive Copyright Agent
Internet Archive
300 Funston Ave.
San Francisco, CA 94118
Email: info@archive.org
Phone: 415-561-6767

How do I know if I can use this?

The Internet Archive does not make guarantees as to the copyright status of items on archive.org and cannot guarantee information posted on item details or collection pages regarding copyright or other intellectual property rights. Our Terms of Use require that users make use of the Internet Archive’s Collections at their own risk and ensure that such use is non-infringing and in accordance with all applicable laws.

The person who uploads an item often provides information related to use rights, either by way of directly entering it in the description field or by selection of a Creative Commons license. The latter, if included by the uploader, will be viewable via a  Creative Commons  logo on the details page, which serves as a link to a description of the specific type of license that the uploader has assigned.

One way to attempt to contact an uploader about information that they have posted is to post a review to the item.

You may also find these resources helpful:
CreativeCommons.org
Lumen Database
Electronic Frontier Foundation 

Please see also:

Who owns the rights to these movies?

Please see the Who owns the rights to the movies that have been uploaded? section at https://help.archive.org/help/movies-and-videos-a-basic-guide/

Are there restrictions on the use of the Prelinger Films?

Please see the Prelinger Archive page at https://help.archive.org/help/prelinger-archive/

Can I search Archive.org by Creative Commons License?

Please see the Can I search by Creative Commons license? section at https://archivesupport.zendesk.com/knowledge/articles/360018359991/en-us

What is non-Commercial Use?

For cultural materials that, broadly defined, belong in a library, the Internet Archive offers free storage, and free bandwidth, forever, for free. As a result, there are now millions of works available through the Archive and most are available only for “non commercial use” and “with attribution.” Sometimes creators choose a Creative Commons license (creativecommons.org) to express this. 

But what does “non-commercial use” mean? We are looking to understand people’s intent, which may be reflected in law in the future. If, collectively, we arrive at a good definition, then we hope many more people will make their works broadly available. This is a start of a definition that we could feel comfortable with. Please let us know what you think via the forum at http://www.archive.org/iathreads/post-view.php?id=111590 . 

As a starting point, the Grateful Dead, a granddaddy of music sharing (called tape trading back in the day) says this: 
• “No commercial gain may be sought by websites offering digital files of our music, 
whether through advertising, exploiting databases compiled from their traffic, or any other means. 
• All participants in such digital exchange acknowledge and respect the copyrights of the performers, writers and publishers of the music. 
•This notice should be clearly posted on all sites engaged in this activity.” 
http://web.archive.org/web/20051130025025/http://www.dead.net/hotline_info/NEW_DOCUMENTS/mp3.html 

How can we interpret this? 

This statement revolves around the website that offers the files, not the company and not the webpage. Therefore, a commercial company could store and offer files non-commercially if they are on a separate website that is isolated from their other sites which might be deemed commercial. We interpret this as the originating host and those sites that present the files to users. 

To broaden this definition, we could substitute “website” to “service” to cover a broader range of methods to find and retrieve information. 

Subscription services: if only subscribers can get to the material, then that is deemed commercial use, even if this is done by a non-profit organization. If non-subscribers can also get equal access, then it is not commercial. 

Ad supported websites that give away materials would be deemed commercial use. Therefore files offered from the AOL.com or NYTimes.com websites would be commercial use. 

Catalog and navigation services, such as search engines, that help direct people to the non-commercial materials on non-commercial sites, we believe, do not violate non-commercial status, but navigation starts to look like re-hosting the materials (therefore possibly commercial use) when it is over a 2 line textual description, as common on Yahoo’s and Google’s web search result pages. There are also deep questions that arise when commercial companies, such as those that operate navigation services, keep and use the files pure navigation. 

We do not believe that this agreement precludes uses that are “fair” (under Title 17, Section 107, of the U.S. Code). Although the agreement allows many uses that would not necessarily be fair uses, such as wholesale copying and rehosting elsewhere, fair use also allows many uses that this agreement does not explicitly mention. 
For example, nytimes.com might be allowed to use a section of the work for news reporting purposes under fair use or other similar principles, even though that use is commercial (because nytimes.com is a for-profit service that makes money selling ads and subscriptions). Some people say that fair use does not exist in the digital world, because licenses and DRM explicitly strip away the power to make fair uses. We disagree: we see declarations that specify non-commercial uses as opening up both non-commercial and fair uses. 

Give-away promotion: if the materials are being given away as part of a promotion, then that is deemed commercial use. For instance if Apple, through its iTunes service gave away material to promote its brand and attract paying customers of other services, that is commercial. If as part of the freely distributed iTunes desktop software, materials were also cataloged and distributed, that would be deemed non-commercial use. 

So what are some non-commercial websites or services? Here are a few: 
www.loc.gov by the Library of Congress 
www.ibiblio.org by the University of North Carolina 
www.archive.org by the Internet Archive 
www.eff.org by the Electronic Frontier Foundation 
web.mit.edu by the Massachusetts Institute of Technology 

A link the Terms of Use for Archive.org is at the bottom of each page. 

How can I contact the person / group who uploaded an item?

Internet Archive is unable to release any contact information for patrons. However, it may be worth your while to post a review for the item in question – this automatically contacts the uploader’s account, notifying them that their upload has been reviewed. You could pose queries/requests for information therein.

]]>