Archive.org – Internet Archive Help Center https://help.archive.org How can we help you? Fri, 26 Jun 2026 21:40:54 +0000 en-US hourly 1 https://wordpress.org/?v=6.8 https://help.archive.org/files/2024/03/cropped-Internet-Archive-Logo-White-on-Black-520x520-1-32x32.png Archive.org – Internet Archive Help Center https://help.archive.org 32 32 FAQ: Publishers Blocking the Wayback Machine https://help.archive.org/help/faq-publishers-blocking-the-wayback-machine/ Fri, 15 May 2026 18:53:26 +0000 https://help.archive.org/?p=1968 Nieman Lab first reported that some publishers and news organizations have begun blocking the Internet Archive’s Wayback Machine from preserving and providing access to archived versions of their websites. 

Since then, journalists, digital rights advocates, historians, librarians, and researchers have raised alarms about the long-term consequences of limiting web preservation.

What is the Wayback Machine?

The Internet Archive is a nonprofit research library with a mission of providing Universal Access to All Knowledge. The Wayback Machine is a service of the Internet Archive that allows people to visit archived versions of web pages. The Wayback Machine has been archiving the web since 1996, helping preserve the historical record of the internet. Learn more about the Wayback Machine.

Why are publishers blocking the Wayback Machine?

As reported by journalists Andrew Deck and Hanaa’ Tameez in Nieman Lab, publishers say they are concerned about AI scraping and unauthorized reuse of their content by generative AI companies.

Some organizations have responded by broadly blocking web archiving systems like the Wayback Machine out of concern that archived material could be accessed or reused by AI systems.

Nicholas Thompson, CEO of The Atlantic, published a video explaining why The Atlantic has blocked the Wayback Machine from preserving its web content, saying:

“[The Wayback Machine] is an amazing resource…But, on the other hand, if you put everything up there, it will all not only go into the big AI companies, but to everybody who’s building some kind of subsidiary AI.”

But those concerns are unfounded, explains journalist Andrew Deck in an interview with Marketplace Tech:

“I think it’s important to say that in our conversations with news publishers, a lot of them were taking this action preemptively out of a fear of proxy scraping rather than direct evidence that it has happened to them already. None of the publishers were able to point to a particular AI company or other kinds of direct evidence that their content had already been scraped by the Wayback Machine.”

Michael Nelson, a computer scientist at Old Dominion University, has described the Wayback Machine and web archiving as “collateral damage” as content owners restrict access over AI scraping concerns. 

That sentiment was echoed by Mark Graham, director of the Wayback Machine, on the Future Knowledge podcast episode, “Preserving the Web in the Age of AI”:

“The Wayback Machine is collateral damage caught up in the conflict between AI companies and publishers.”

Who uses the Wayback Machine?

The Wayback Machine is widely used by journalists, fact-checkers, researchers, lawyers, courts, librarians, academics, and even publishers themselves.

More than 100 news articles every month reference, cite, or rely on material preserved by the Wayback Machine to verify claims, recover deleted information, or provide historical context.

In a Future Knowledge podcast interview, Mark Graham recalled a conversation at The New York Times:

“A senior researcher came up to me and said, ‘Oh my God, Mark, thank you so much for the Wayback Machine. We use you all the time. There is material available that we’ve used from the Wayback Machine that we can’t even find in our own archives.’”

Journalist Rachel Maddow also publicly defended the archive:

“The Internet Archive is a national treasure. I use it daily, and have for many, many years. I cannot imagine doing the work I do without it.”

What are critics saying about the blocking?

A broad coalition of journalists, digital rights advocates, and internet historians has warned that blocking the Wayback Machine—and web archiving tools like it—could have serious long-term consequences.

In Techdirt, Mike Masnick argued that publishers may regret these blocking decisions because they undermine preservation of the public record.

Joe Mullin, senior policy analyst at the Electronic Frontier Foundation, similarly warned that blocking the Internet Archive “won’t stop AI” but could erase important historical records from the web.

Meanwhile, coverage by journalist Kate Knibbs in Wired brought broader public attention to the issue and the risks facing digital preservation infrastructure.

Have journalists spoken out in support of the Wayback Machine?

Yes.

More than 200 journalists signed a public statement applauding the Internet Archive’s role in preserving the public record.

In response, Mark Graham published a public thank-you letter:

“Your support for the Wayback Machine sends a clear message: preserving the record matters.”

Does the Internet Archive respect publisher concerns?

Yes.

The Internet Archive works with publishers and rights holders to balance preservation, access, and responsible stewardship of digital materials.

What about AI scraping?

As Mark Graham explained in Techdirt:

“The Wayback Machine is built for human readers. We use rate limiting, filtering, and monitoring to prevent abusive access, and we watch for and actively respond to new scraping patterns as they emerge.”

What’s at stake when publishers block the Wayback Machine?

Every day that preservation systems are blocked leaves holes in the public record of the web. If preservation systems are weakened or blocked at scale, future generations will lose access to major parts of our digital history.

The size of the problem is significant. A 2024 study by Pew Research Center found that 38% of webpages from 2013 were no longer accessible a decade later, with roughly a quarter of pages sampled across the decade disappearing entirely. But loss on the web is not inevitable. New analysis by Internet Archive data scientist Sawood Alam found that the Wayback Machine has preserved roughly 15% of those otherwise vanished pages, saving reporting, citations, and pieces of the historical record that would no longer exist online.

As Mike Masnick wrote in Techdirt: “Blocking the Internet Archive isn’t going to stop AI training. What it will do is ensure that significant chunks of our journalistic record and historical cultural context simply… disappear.”

And as Mark Graham wrote:

“Preserving the public record is not optional. It is essential infrastructure for a functioning democracy.”


Chronological Reading List of Key Articles


FIRST PUBLISHED: May 15, 2026 CDF
LAST UPDATED: June 26, 2026 CDF

]]>
Does the Internet Archive Have an Onion Address? https://help.archive.org/help/does-the-internet-archive-have-an-onion-address/ Wed, 11 Mar 2026 15:52:29 +0000 https://help.archive.org/?p=1850 Does the Internet Archive have an onion address?

Yes, the Internet Archive has an onion address. The Internet Archive can be accessed via the Tor network at its onion address: archivep75mbjunhxc6x4j5mwjmomyxb573v42baldlqu56ruil2oiad.onion

What is an onion address?

Tor (The Onion Router) is a privacy-focused network that helps protect users’ identities and browsing activity by routing traffic through encrypted layers. Visiting the Internet Archive through Tor allows users to explore the Wayback Machine, books, audio, video, and other collections with an added layer of anonymity, which is an important option for researchers, journalists, and anyone seeking greater privacy or access in regions where the open web may be restricted.

]]>
Federal Depository Library Program (FDLP) FAQ https://help.archive.org/help/federal-depository-library-program-fdlp-faq/ Tue, 29 Jul 2025 02:38:39 +0000 https://help.archive.org/?p=1757 Internet Archive has been designated a federal depository library. Here are some of the frequently asked questions about the designation:

  1. Why did Internet Archive pursue becoming a federal depository library?
    • This is the culmination of years of work building Democracy’s Library, a free, open, online collection of government research and publications from around the world. As Brewster Kahle, digital librarian of the Internet Archive, explained to KQED:

      “By being part of the program itself, it just gets us closer to the source of where the materials are coming from, so that it’s more reliably delivered to the Internet Archive, to then be made available to the patrons of the Internet Archive or partner libraries.”
  2. What does the FDLP designation mean for Internet Archive in practical terms?
    • It means that Internet Archive will get more documents from the Government Publishing Office and from other FDLP partner libraries that we will preserve, digitize (if in physical form), and make available through archive.org. Internet Archive currently receives donations of government documents from libraries, but there are some materials that can only be donated to other FDLP libraries. By being part of the program, Internet Archive can now work with those libraries, and GPO directly, to preserve, digitize, and provide access to even more public information.
  3. How are FDLP designations made?
    • Libraries can become part of the Federal Depository Library Program (FDLP) in two main ways. First, some libraries are designated by law—for example, land grant universities automatically qualify. Second, members of Congress can designate libraries in their district or state to join the FDLP, as long as there’s a vacancy and the library can show it’s capable of serving the public’s need for access to government information. The Internet Archive was designated as a federal depository library on July 24, 2025, by Senator Alex Padilla of California.
  4. Does this mean the government owns the Wayback Machine?
    • No. Becoming a Federal Depository Library does not give the government any ownership of or control over the Wayback Machine or the Internet Archive’s collections. The Federal Depository Library Program (FDLP) is a voluntary partnership that helps libraries provide free public access to government publications. The Internet Archive remains an independent, nonprofit research library, and continues to curate and operate its collections—including the Wayback Machine—on its own terms.
  5. What does this designation mean for Internet Archive?
    • Becoming a member of the Federal Depository Library Program furthers Internet Archive’s mission to provide “Universal access to all knowledge.” It is the culmination of years of existing collaborations and partnerships with FDLP libraries. With our new designation as a federal depository library, we’re officially part of the national network that provides the public with free access to government publications. That means we can now acquire, preserve, and share even more official government documents—further expanding our collections and strengthening public access to government information.
]]>
Search – Building Powerful, Complex Queries https://help.archive.org/help/search-building-powerful-complex-queries/ Thu, 04 Apr 2024 17:37:45 +0000 https://help.archive.org/?p=1629 The Internet Archive has tens of millions of items, so sometimes finding exactly what you’re looking for can be difficult. But if you learn to build more complex search queries, you can narrow your results much more quickly. This video (or the advanced search page) is a good place to start.

In general, using Boolean Operators and doing fielded searches are your best bets.

Boolean Operators

The video above contains a quick explanation, but you can also read more about Boolean Operators elsewhere on the Internet. They are useful on archive.org, but you can use them in other search engines too. As a quick recap, the most common are:

  • AND – narrows your search
  • OR – widens your search
  • AND NOT – excludes things from your search
  • (  ) – parentheses can be used to group search terms
  • ”  “ – double quotes will only return searches with an exact match for that phrase within
  • [ ]  – square brackets can be used for ranges when ranges are allowed

Fielded Searches

If you would like to explore doing more fielded searches, use the Internet Archive metadata schema to figure out what fields you can search and what their values tend to look like. Here are some common fields you might want to search:

  • title – the title of the work
    • e.g. title:”war and peace” will find items with the exact phrase “war and peace” in the title field
  • subject – terms that describe the content of the work (also referred to as topics or keywords)
    • e.g. subject:mythology will find items with the word mythology in their subject fields
  • creator – the person (or entity) that created the work, e.g. the author, director, conductor, etc.
    • e.g. creator:(Dickens AND Charles) will find all items with both Dickens and Charles in the creator field
  • date – the date the work was published, often in YYYY or YYYY-MM-DD format
    • e.g. date:1922 will find all items with a publish date within the year 1922
]]>
Why are so many books listed as “Borrow Unavailable” at the Internet Archive https://help.archive.org/help/why-are-so-many-books-listed-as-borrow-unavailable-at-the-internet-archive/ Sat, 16 Mar 2024 05:24:48 +0000 https://help.archive.org/?p=1618 Summary

Books that are shown as “Borrow Unavailable” mean they cannot be borrowed by our patrons, including books you may have previously read or consulted. 

In 2020, our library was sued by four of the world’s largest publishers—Hachette, HarperCollins, Penguin Random House, and Wiley—for lending books via the library practice known as controlled digital lending. That lawsuit, Hachette v. Internet Archive, is currently on appeal after the lower court (the United States District Court for the Southern District of New York) heard the case and ruled in favor of the publishers. In this decision against our library, Judge Koeltl issued an injunction that limits what we can do with our digitized books—namely, we can no longer lend those books to our patrons. The injunction does not affect our accessibility program—the removed books are still available to patrons with print disabilities.

Additionally, the Association of American Publishers (AAP), the trade organization behind the lawsuit, worked with some of its member publishers (listed below) that were not named in the lawsuit to demand that we remove their books from our library.

As a result, more than 500,000 books in our collection are not currently available for borrowing, including more than 1,300 banned and challenged books. We understand that this is a devastating loss for our patrons. Fortunately, other countries and international library organizations are moving to support controlled digital lending. For inquiries, please contact patron services at info@archive.org.

Plaintiff Publishers

  • Hachette Book Group
  • HarperCollins
  • Penguin Random House
  • Wiley

Other Publishers Coordinated by the Association of American Publishers (AAP)

  • American Chemical Society
  • American Reading Company
  • BiggerPockets Publishing
  • Bloomsbury
  • Bookpress Publishing
  • Cambridge University Press
  • Chronicle Books
  • De Gruyter
  • Elsevier
  • Fordham University Press
  • Getty Publications
  • Hansen Publishing Group
  • Harvard University Press
  • Holiday House-Peachtree-Pixel+Ink
  • Imbrifex Books
  • Lynne Rienner Publishers
  • Macmillan
  • MedMaster
  • Melville House Publishing
  • Moody Publishers
  • Pearson
  • Princeton University Press
  • Sage Publications
  • Scholastic
  • Simon & Schuster
  • Springer
  • Taylor & Francis
  • Teacher Created Materials
  • The American University in Cairo Press
  • University of California Press
  • University of Chicago Press
  • University of Massachusetts Press
  • University of Minnesota Press
  • University of Texas Press
  • University of Wisconsin Press
  • University Press of Colorado
  • Valancourt Books
  • W. W. Norton & Company
  • Wesleyan University Press
  • Wolters Kluwer
  • Yale University Press
]]>
Lists- A basic guide https://help.archive.org/help/lists-a-basic-guide/ Fri, 15 Mar 2024 23:09:38 +0000 https://help.archive.org/?p=1551 Lists are a convenient way to organize content on the Archive. They can be created by any logged-in user. A list can have a title and description, as well as be configured to be public or private. A private list can only be seen by the patron who created it. Public lists can be seen by anyone and can also be shared.

Adding items to a list

  1. Go to the item’s detail page.
  2. Click or tap the “Add to list” button underneath the top theater section.
  3. Choose a list to add the item to by clicking or tapping on it. To create a new list and add the item to it, click or tap “Create new list” and enter the list information.

Removing items from a list

An item can be removed from a list either from the item’s detail page or from the “My Lists” page.

From the detail page

  1. Click or tap the “Add to list” button.
  2. Find the checkmarked list that the item is to be removed from, and click or tap on that list.

From the “My Lists” page

  1. Go to your Profile Page > My Lists.
  2. Select the list in the left column from which the item is to be removed.
  3. When the list contents are displayed, click or tap “Remove items…” in the top right list of actions.
  4. Select the item or items to be removed from that list.
  5. Click or tap “Okay”.

Managing lists

To manage your lists, go to your Profile Page > My Lists. There you can:

  • View all your lists
  • Edit the information for any of your lists
  • Create a new list
  • Delete a list
  • Remove items from a list

Creating a new list

  1. Click or tap “Add list” in the left column.
  2. Enter a name for the list and optionally, a description. If you would like to prevent others from seeing your list, mark it as Private.
  3. Click or tap “Save”.

You can now add items to this list from any item detail page.

Editing a list

  1. Select a list to edit by clicking or tapping on the list name in the left column.
  2. Click or tap the pencil icon in the main pane to the right of the list name.
  3. Edit the list properties and click “Save”.

Deleting a list

  1. Select a list to edit by clicking or tapping on the list name in the left column.
  2. Click or tap the trash can icon at the edge of the screen to the right of the list name.
  3. Click “Delete” to confirm the deletion.

Removing items from a list

  1. Select a list to edit by clicking or tapping on the list name in the left column.
  2. Click or tap “Remove items” from the list of actions at the top right.
  3. Select the items to delete by clicking or tapping on their tiles. A checkmark will be displayed in the top right corner of the tile for each item to be deleted.
  4. After all desired tiles have been selected, click or tap “Remove selected items”.
  5. On the subsequent confirmation dialog box, click or tap “Remove items” to confirm the deletion.
]]>
Files, Formats, and Derivatives – file definitions https://help.archive.org/help/files-formats-and-derivatives-file-definitions-2/ Fri, 15 Mar 2024 22:46:32 +0000 https://help.archive.org/?p=896
formatextensionmediatypereview
Metadatafiles.xmlallthe manifest that records all of the files available for this book; also gives 2 checksums and a format definition for each file; provides the only mechanism for validating that the component data has been downloaded successfully
Simple File Verification.sfvallSimple file verification (SFV) is a file format for storing CRC32 checksums of files to verify the integrity of files. SFV is used to verify that a file has not been corrupted, but it does not otherwise verify the file’s authenticity. https://en.wikipedia.org/wiki/Simple_file_verification
Metadatameta.xmlallInternet Archive’s internal “management” metadata; a proprietary XML format, this file includes information about the scan event (date, # of pages, operator, station, etc.), the contributor, basic bib data (title, author, subject, language), and a set of identifiers
Windows Media Audio.wmaaudioWindows Media Audio (WMA) is a series of audio codecs and their corresponding audio coding formats developed by Microsoft. https://en.wikipedia.org/wiki/Windows_Media_Audio
WAVE.wavaudioWaveform Audio File Format (WAV) is an audio file format standard for storing an audio bitstream. https://en.wikipedia.org/wiki/WAV
Ogg Vorbis.oggaudioVorbis is a free and open-source software project. The project produces an audio coding format and software reference encoder/decoder (codec) for lossy audio compression. Vorbis is most commonly used in conjunction with the Ogg container format and it is therefore often referred to as Ogg Vorbis. https://en.wikipedia.org/wiki/Vorbis
VBR MP3.mp3audioVBR (Variable Bitrate) MP3 is a coding format for digital audio https://en.wikipedia.org/wiki/MP3
VBR M3U.m3uaudioVBR (Variable Bitrate) M3U is a computer file format for a multimedia playlist. One common use of the M3U file format is creating a single-entry playlist file pointing to a stream on the Internet. The created file provides easy access to that stream and is often used in downloads from a website, for emailing, and for listening to Internet radio. https://en.wikipedia.org/wiki/M3U
Shorten.shnaudioShorten (SHN) is a file format used for compressing audio data. It is a form of data compression of files and is used to losslessly compress CD-quality audio files. https://en.wikipedia.org/wiki/Shorten_(codec)
MP3.mp3audioMP3 is a coding format for digital audio https://en.wikipedia.org/wiki/MP3
M3U.m3uaudioM3U is a computer file format for a multimedia playlist. One common use of the M3U file format is creating a single-entry playlist file pointing to a stream on the Internet. The created file provides easy access to that stream and is often used in downloads from a website, for emailing, and for listening to Internet radio. https://en.wikipedia.org/wiki/M3U
MP3 Samplesample.mp3audioxLimited length MP3 audio file derived from source audio file. Typically 30 seconds in length.
Flac.flacaudioFLAC is an audio coding format for lossless compression of digital audio, developed by the Xiph.Org Foundation, and is also the name of the free software project producing the FLAC tools, the reference software package that includes a codec implementation. https://en.wikipedia.org/wiki/FLAC
AIFF.aiffaudioAudio Interchange File Format (AIFF) is an audio file format standard used for storing sound data for personal computers and other electronic audio devices. https://en.wikipedia.org/wiki/Audio_Interchange_File_Format
Advanced Audio Coding.m4aaudioAdvanced Audio Coding (AAC) is an audio coding standard for lossy digital audio compression. Designed to be the successor of the MP3 format, AAC generally achieves higher sound quality than MP3 encoders at the same bit rate. https://en.wikipedia.org/wiki/Advanced_Audio_Coding
Spectrogramspectrogram.pngaudioxA visual representation of the spectrum of frequencies of a signal as it varies with time.
Columbia Fingerprint.afpkaudiox“audio fingerprinting” to enable comparing audio tracks together for “the same” tracks or portions of them
Columbia Fingerprintffp.txtaudiox“audio fingerprinting” to enable comparing audio tracks together for “the same” tracks or portions of them
Essentia High GZesshigh.json.gzaudioxhistorical audio format that tried to do analysis like beats-per-minute, deductions of “genre” of music, etc.
Essentia Low GZesslow.json.gzaudioxhistorical audio format that tried to do analysis like beats-per-minute, deductions of “genre” of music, etc.
Flac FingerPrint.ffpaudioxa community-specific checksum for flac files, important to etree community
ZIP.zipdataZIP is an archive file format that supports lossless data compression. A ZIP file may contain one or more files or directories that may have been compressed. https://en.wikipedia.org/wiki/ZIP_(file_format)
Rich Text Format.rtfdataThe Rich Text Format (often abbreviated RTF) is a proprietary document file format. Most word processors are able to read and write some versions of RTF. https://en.wikipedia.org/wiki/Rich_Text_Format
OpenDocument Text Document.odtdataThe Open Document Format for Office Applications (ODF), also known as OpenDocument, is an open standard file format for spreadsheets, charts, presentations and word processing documents using ZIP-compressed XML files https://en.wikipedia.org/wiki/OpenDocument
HTML.htmldataThe HyperText Markup Language or HTML is the standard markup language for documents designed to be displayed in a web browser. It can be assisted by technologies such as Cascading Style Sheets (CSS) and scripting languages such as JavaScript. https://en.wikipedia.org/wiki/HTML
Shockwave.swfdataSWF is an Adobe Flash file format used for multimedia, vector graphics. SWF files can contain animations or applets of varying degrees of interactivity and function. They may also occur in programs, commonly browser games, using ActionScript. https://en.wikipedia.org/wiki/SWF
RAR.rardataRAR is a proprietary archive file format that supports data compression, error recovery and file spanning. https://en.wikipedia.org/wiki/RAR_(file_format)
OpenType Font.otfdataOpenType is a format for scalable computer fonts.. https://en.wikipedia.org/wiki/OpenType
MIDI.middataMIDI is a technical standard that describes a communications protocol, digital interface, and electrical connectors that connect a wide variety of electronic musical instruments, computers, and related audio devices for playing, editing, and recording music. https://en.wikipedia.org/wiki/MIDI
Word Document.docdataMicrosoft Word is a word processing software developed by Microsoft. https://en.wikipedia.org/wiki/Microsoft_Word
Powerpoint.pptdataMicrosoft PowerPoint is a presentation program. PowerPoint was originally designed to provide visuals for group presentations within business organizations, but has come to be very widely used in many other communication situations, both in business and beyond. https://en.wikipedia.org/wiki/Microsoft_PowerPoint
Excel.xlsdataMicrosoft Excel is a spreadsheet developed by Microsoft. https://en.wikipedia.org/wiki/Microsoft_Excel
JSON.jsondataJSON is an open standard file format and data interchange format that uses human-readable text to store and transmit data objects consisting of attribute–value pairs and arrays (or other serializable values). https://en.wikipedia.org/wiki/JSON
TAR.tardataIn computing, tar is a computer software utility for collecting many files into one archive file, often referred to as a tarball, for distribution or backup purposes. https://en.wikipedia.org/wiki/Tar_(computing)
Text.txtdataIn computing, plain text is a loose term for data (e.g. file contents) that represent only characters of readable material but not its graphical representation nor other objects (floating-point numbers, images, etc.). https://en.wikipedia.org/wiki/Plain_text
GZIP.gzdatagzip is a file format and a software application used for file compression and decompression. https://en.wikipedia.org/wiki/Gzip
Flash Video.flvdataFlash Video is a container file format used to deliver digital video content (e.g., TV shows, movies, etc.) over the Internet using Adobe Flash Player version 6 and newer. https://en.wikipedia.org/wiki/Flash_Video
Cascading Style Sheet.cssdataCascading Style Sheets (CSS) is a style sheet language used for describing the presentation of a document written in a markup language such as HTML. https://en.wikipedia.org/wiki/CSS
ISO Image.isodataAn optical disc image (or ISO image, from the ISO 9660 file system used with CD-ROM media) is a disk image that contains everything that would be written to an optical disc, disk sector by disc sector, including the optical disc file system. https://en.wikipedia.org/wiki/Optical_disc_image
Adobe Illustrator.aidataAdobe Illustrator Artwork (AI) is a proprietary file format developed by Adobe Systems for representing single-page vector-based drawings in either the EPS or PDF formats. The .ai filename extension is used by Adobe Illustrator. https://en.wikipedia.org/wiki/Adobe_Illustrator_Artwork
Tab-Separated Values.tsvdataA tab-separated values (TSV) file is a simple text format for storing data in a tabular structure, e.g., a database table or spreadsheet data, and a way of exchanging information between databases. https://en.wikipedia.org/wiki/Tab-separated_values
7Z.7zdata7z is a compressed archive file format that supports several different data compression, encryption and pre-processing algorithms. https://en.wikipedia.org/wiki/7z
Windows Executable.exedata.exe is a common filename extension denoting an executable file (the main execution point of a computer program) for Microsoft Windows. https://en.wikipedia.org/wiki/.exe
Animated GIF.gifimageThe Graphics Interchange Format is a bitmap image format that was developed by a team at the online services provider CompuServe. https://en.wikipedia.org/wiki/GIF
TIFF.tiffimageTag Image File Format, abbreviated TIFF or TIF, is an image file format for storing raster graphics images, popular among graphic artists, the publishing industry, and photographers. https://en.wikipedia.org/wiki/TIFF
PNG.pngimagePortable Network Graphics is a raster-graphics file format that supports lossless data compression. https://en.wikipedia.org/wiki/Portable_Network_Graphics
JPEG.jpgimageJPEG is a commonly used method of lossy compression for digital images, particularly for those images produced by digital photography. https://en.wikipedia.org/wiki/JPEG
JPEG 2000.jp2imageJPEG 2000 (JP2) is an image compression standard and coding system. https://en.wikipedia.org/wiki/JPEG_2000
Web Video Text Tracks.vttmoviesWebVTT (Web Video Text Tracks) is a (W3C standard for displaying timed text in connection with the HTML5 <track> element. https://en.wikipedia.org/wiki/WebVTT
WebM.webmmoviesWebM is an audiovisual media file format. It is primarily intended to offer a royalty-free alternative to use in the HTML5 video and the HTML5 audio elements. https://en.wikipedia.org/wiki/WebM
Ogg Video.ogvmoviesTheora is a free lossy video compression format. It is is most commonly used in conjunction with the Ogg container format. https://en.wikipedia.org/wiki/Theora
Checksums.md5moviesThe MD5 message-digest algorithm is a cryptographically broken but still widely used hash function producing a 128-bit hash value. https://en.wikipedia.org/wiki/MD5
Matroska.mkvmoviesThe Matroska Multimedia Container is a free and open container format, a file format that can hold an unlimited number of video, audio, picture, or subtitle tracks in one file. https://en.wikipedia.org/wiki/Matroska
MPEG4.m4vmoviesThe M4V file format is a video container format developed by Apple and is very similar to the MP4 format. The primary difference is that M4V files may optionally be protected by DRM copy protection. https://en.wikipedia.org/wiki/M4V
QuickTime.movmoviesQuickTime is a video format that is particularly suited for editing, as it is capable of importing and editing in place (without data copying). https://en.wikipedia.org/wiki/QuickTime_File_Format
MPEG4.mpeg4moviesMPEG-4 is a method of defining compression of visual (AV) digital data. https://en.wikipedia.org/wiki/MPEG-4
MPEG2.mpegmoviesMPEG-2 is a standard for “the generic coding of moving pictures and associated audio information”. https://en.wikipedia.org/wiki/MPEG-2
MPEG2.mpgmoviesMPEG-2 is a standard for “the generic coding of moving pictures and associated audio information”. https://en.wikipedia.org/wiki/MPEG-2
512Kb MPEG4512kb.mp4moviesxLow resolution MPEG4 video file
Thumbnailthumb.jpgmoviesxImages of video captured approximated every 30 seconds. They are used in the player scrubber
h.264 IAia.mp4moviesxDerived h.264 file intended to create web-friendly version of uploaded source mp4 that does not meet the minimum criteria for optimal use in the online media player.
Closed Caption Textcc5.txtmoviesxClosed captions text file captured with tv archive recordings
SubRipalign.srtmoviesxClosed Captions in TV Archive items adjusted to better align with the AV
SubRipcc5.srtmoviesxClosed Captions in TV Archive items
Cinepack.avimoviesCinepak is a lossy video codec developed by Peter Barrett at SuperMac Technologies, and released in 1991 with the Video Spigot, and then in 1992 as part of Apple Computer’s QuickTime video suite. https://en.wikipedia.org/wiki/Cinepak
ASRasr.jsmoviesxAutomatic Speech Recognition closed captions. Computer generated from mp3 audio files that are converted to text files.
ASRasr.srtmoviesxAutomatic Speech Recognition closed captions formatted to run in conjucntion with the related video file. Computer generated from mp3 audio files that are converted to text files.
h.264.mp4moviesAdvanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
Windows Media.wmvmoviesAdvanced Systems Format (wmv) is Microsoft’s proprietary digital audio/digital video container format, especially meant for streaming media. https://en.wikipedia.org/wiki/Advanced_Systems_Format
h.264h.264 720Pmoviesx720px1080p h.264 file. Advanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
h.264h.264 HDmoviesx720px1080p h.264 file. Advanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
for tvarchive.xmlmoviesxTV Archive minimal metadata to create full metadata for a show (eg: program title & description, scheduled duration, etc.)
h.264h.264 popcornmoviesxOnline directly in-the-browser user edited audio/video editor files that will playback arbitrary audio & video files, add textual overlays, maps, and more as well
JPEG Thumbthumb.jpgmoviesxA smaller version of various item image files
JSONalign.jsonmoviesxCaptions alignment (audio wave form vs. captions) to reduce the “drift” between what is spoken vs. what got captioned. They can often have 2-10 seconds of distance between displayed words/captions and heard audio
Derivation Rulesrules.confmovies/audioxPrevents lossy derivatives of source data files in audio and video items
Android Package Archive.apksoftwareThe Android Package with the file extension apk is the file format used by the Android operating system, and a number of other Android-based operating systems for distribution and installation of mobile apps, mobile games and middleware. https://en.wikipedia.org/wiki/Apk_(file_format)
Emulator Screenshotscreenshot.pngsoftwarexScreen capture of an emulated computer game
Mac OS X Disk Image.dmgsoftwareApple Disk Image is a disk image format commonly used by the macOS operating system. When opened, an Apple Disk Image is mounted as a volume within the Finder. https://en.wikipedia.org/wiki/Apple_Disk_Image
iOS App Store Package.ipasoftwareAn .ipa (iOS App Store Package) file is an iOS application archive file which stores an iOS app. https://en.wikipedia.org/wiki/.ipa
Amiga Disk File.adfsoftwareAmiga Disk File (ADF) is a file format used by Amiga computers and emulators to store images of floppy disks. https://en.wikipedia.org/wiki/Amiga_Disk_File
Windows Screensaver.scrsoftwareA screensaver is a computer program that blanks the display screen or fills it with moving images or patterns, when the computer has been idle for a designated time. https://en.wikipedia.org/wiki/Screensaver
Log.logtextsxThere are several logs from scanning, republishing, etc. e.g. Cloth Cover Detection Log, various Republisher Logs, and then the plan Log format for Scribe logs.
PDF.pdftextsThe presentation version on BHL in PDF format. Low quality; sufficient for printing and reading text
Metadatareviews.xmltextsxThe meta.xml file contains all of the item-level metadata for reviews
Metadatameta.xmltextsxThe meta.xml file contains all of the item-level metadata for an item (e.g. title, description, creator, etc.).
MARC Binarymarc.xmltextsthe MARC (bibliographic description) data in XML. MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
MARC Binarymeta.mrctextsthe binary MARC record as retrieved using z39.50. MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
Single Page Original JP2 Tarorig_jp2.tartextsSome books are so large that the volume of images exceed the maximum size for a ZIP archive. For these books, the images are compressed and delivered using TAR. These TAR archives average 2.07 gb and occur .39% of the time (738 out of 191,568 books total).
High quality; Best for use and printing of plates, illustrations, detailed figures and tables
DjVu.djvutextsSimilar to PDF, a proprietary compressed document format.
Low quality; sufficient for printing and reading text
Scandatascandata.xmltextsxScandata is an XML file containing specific per-image information, including if the image should be included in any of the produced formats. The module will find, parse and honors these files if they exist.
Text PDF.pdftextsxPortable Document Format files, containing MRC-compressed images and the OCR result as a hidden (selectable, searchable) text layer. (In some cases, the PDF files can have a slightly different suffix, but the extension remains .pdf)
Item Imageitemimage.pngtextsxPNG image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
chOCRchocr.html.gztextsxOCR results with character-level granularity
Dublin Coredc.xmltextsOAI record in Dublin Core (bibliographic description) XML. Dublin Core is a set of metadata elements that provide a small and fundamental group of text elements through which most resources can be described and cataloged; a metadata format for describing resources.
Metadatameta.sqlitetextsxMetadata for file sync via an sqlite database
Name Metadatanames.xmltextslist, by page, of all the scientific names found in the book; presented in xml format
Item Imageitemimage.jpgtextsxJPG image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
Item Tile__ia_thumb.jpgtextsxItem thumbnail image used in search results tiles
Abbyy ZIPabbyy.gztextsGZipped version of the full ABBYY FineReader XML output, which includes all character-level information (confidence, location, etc.)
Item Imageitemimage.giftextsxGIF image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
EPUB.epubtextsEPUB is an e-book file format that uses the “.epub” file extension. The term is short for electronic publication and is sometimes styled ePub. EPUB is supported by many e-readers, and compatible software is available for most smartphones, tablets, and computers.
DAISYtextsDigital accessible information system (DAISY) is a technical standard for digital audiobooks, periodicals, and computerized text. DAISY is designed to be a complete substitute for print material and is specifically designed for use by people with “print disabilities”, including blindness, impaired vision, and dyslexia. https://en.wikipedia.org/wiki/Digital_Accessible_Information_System
Archive BitTorrentarchive.torrenttextsxDerived torrent file that contains files information on files in an item. archive.org does not seed files.
PNGslip.pngtextsBook scanning slips that get uploaded to reserve an identifier so as not to have to wait hours for a full book to upload
hOCRhocr.htmltextsxBarring any failures in the OCR process, after upload, every item will get one or more hocr.html files which represent the results of OCR jobs. Each hocr.html file contains results for all pages in one set of images (book, PDF, or otherwise), with text, bounding boxes, and confidence at the word level.

For those seeking more detailed OCR results, each _hocr.html file should also have a corresponding chocr.html.gz file, with character-level granularity. (The exact meaning of “character” differs, of course, per script or language).
Generic Raw Book Zipimages.ziptextsxA zip imagestack file formatted to derive the files necessary to create a flip book, pdf and other text formats
Single Page Processed JP2 ZIPjp2.ziptextsA ZIP archive of all of the cleaned, cropped, etc. JP2 page images. These are the highest quality, least modified images that are available after the raw/orig file set.
High quality; Best for use and printing of plates, illustrations, detailed figures and tables
Generic Raw Book Tarjp2.tartextsxA tar imagestack file formatted to derive the files necessary to create a flip book, pdf and other text formats
OCR Page Indexhocr_pageindex.json.gztextsxa simple JSON array annotating where each individual page element starts in the hocr.html file, enabling quick fast-forwarding to an individual page without parsing all the XML.
MARC Sourcemetasource.xmltextsa proprietary XML file recording where the MARC record came from (catalog, operator, zquery, etc.) MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
Metadatascandata.xmltextsa proprietary XML file recording information about each page image (handSide, cropBox, original width & height, etc.)
OCR Search Texthocr_searchtext.txt.gztextsxa plaintext file that is ingested by the full text search engine.
Djvu XMLdjvu.xmltextsa modified version of the DjVu XML standard, these files can also be used to read OCR results, but the recommendation is to instead parse the hOCR files.
Page Numbers JSONpage_numbers.jsontextsxA map of page numbers auto-detected in a book. If the confidence score is high enough, they are sometimes added to scandata.xml
JSONevents.jsontextsxA json file containing information about Republisher events. The format is deprecated and no longer used
DjVUTXTdjvu.txttextsa human-readable plaintext version of the generated djvu.xml file. OCR stands for “Optical Character Recognition;” the conversion of images of text into text characters
Biodiversity Heritage Library METSbhlmets.xmltextsxA format created and used by Biodiversity Heritage Library (BHL)
Comic Book RAR.cbrtextsA comic book archive or comic book reader file (also called sequential image file) is a type of archive file for the purpose of sequential viewing of images, commonly for comic books. https://en.wikipedia.org/wiki/Comic_book_archive
Comic Book ZIP.cbztextsA comic book archive or comic book reader file (also called sequential image file) is a type of archive file for the purpose of sequential viewing of images, commonly for comic books. https://en.wikipedia.org/wiki/Comic_book_archive
Grayscale PDFbw.pdftextsA black and white PDF compiled using binarized versions of the images. The binarized images are not made available.
Low; sufficient for printing with low cost of ink or printing only text images
MOBI.mobitexts.mobi is an e-book file format that is primarily used for Kindle e-readers. https://en.wikipedia.org/wiki/Mobipocket
Web ARChive.warcwebThe Web ARChive (WARC) archive format specifies a method for combining multiple digital resources into an aggregate archive file together with related information. https://en.wikipedia.org/wiki/Web_ARChive
Web ARChive GZwarc.gzwebA compressed Web ARChive (WARC) archive using gzip,a file format and a software application used for file compression and decompression. The format specifies a method for combining multiple digital resources into an aggregate archive file together with related information. https://en.wikipedia.org/wiki/Web_ARChive
]]>
The Internet Arcade https://help.archive.org/help/the-internet-arcade/ Fri, 15 Mar 2024 20:23:13 +0000 https://help.archive.org/?p=1382 What is the Internet Arcade?
The Internet Arcade is a collection of emulated arcade games from the 1970s-1990s that can be played in your browser. It is located here. There are similar collections of playable console games (the Console Living Room) and general computer software (the Software Library).

How is it Playing Arcade Games in my Browser?
The Internet Arcade uses a program called JSMESS, which is a Javascript port of the MESS and MAME emulator projects. MESS/MAME have been developed over nearly 20 years and are able to emulate hundreds of computer systems and thousands of console and arcade games. A volunteer group has been able to convert MESS/MAME into pure Javascript and make it run in most modern browsers.

What Plugins are Needed?
There are no plugins needed to run the Internet Arcade. It uses 100% Javascript (not to be confused with Java), which is a scripting module inside all modern browsers that has great flexibility for running code, playing sound and video, and doing everything necessary to provide an arcade game in a window. Ironically, if the system is not working for you, a plugin may be preventing it: there are a number of plugins, such as NoScript, which automatically turn off Javascript processing for a site and require you to turn it back to run. If that is the case, the Arcade will not function – please enable Javascript on archive.org to run the Arcade.

How do I Play a Game on the Arcade?
In each entry for a game on the Arcade, you are taken to a page with a description of the game, and a screenshot in the right-hand corner of the gameplay. A line underneath the screenshot says “Run an in-browser emulation of the program”. You can click on the screenshot or the word “Run” to go to the Player page. On the Player page, you are shown a box and underneath it controls for Fullscreen, Mute/Unmute, Dark Background, and possibly others. Inside the box, there should be a MAME or MESS logo. Clicking inside this box, or hitting the spacebar, should start a disk icon spinning and the program will load. When the program is finished loading, the disk icon will stop spinning and the box will expand out to the resolution of the given program. At this point, the arcade machine will begin running. If you do not see the MESS/MAME logo, the program will not start. See other FAQ questions for possible solutions to this problem.

I Don’t See Anything in the Box.
If you do not see a MAME/MESS logo in the box above the “Fullscreen, Dark Background, Mute” buttons on the player page, then JSMESS is not running in your browser for some reason. Some possible reasons to investigate:

Are you running a script blocker like NoScript, that blocks Javascript?
Does your browser have Javascript disabled?
JSMESS can take a few seconds to load – wait 30 seconds to see if the logo appears.
JSMESS generally runs in Firefox, Chrome, Opera, IE and Safari. Are you running a different browser than these?
Is your browser a recent version? JSMESS prefers browsers from the last few months (although it should run, albeit poorly, in earlier versions).
Are you low on memory? Disk space?
If none of these seem to apply, contact us with your setup and situation as you see it.
I Don’t Hear Any Sound.

For reasons that we will explain, sound is muted by default on JSMESS. To enable sound, you (currently) need to start a program (i.e., click on the logo), wait for the arcade machine to start, and then hit the “Unmute” button at the bottom of the running game. This will set a cookie for “Unmute” and after you hit Refresh (F5) on your browser, all later games will have sound. We are aware this is clunky, and intend to rewrite our Player to more intuitively work in the future.

The Sound Sounds Horrible/Scratchy/Distorted!
The JSMESS program uses a standard called “Web Audio” that is still in its early stages – as a result, the JSMESS program is extremely burdensome to this standard, and unless your machine is very fast and the arcade game being run a simpler one, the sound can easily distort, even when doing something like switching between tabs or moving the mouse! This is why the program is, by default, muted. As of November, 2014, a new Web Audio specification has been proposed that allows Javascript programs like JSMESS to run audio more dependably, as we expect for sound and video, and the committees in charge of this specification are very aware of JSMESS as a real-world example of how to improve their specification. We currently can only wait, at which point newer versions of browsers will have much better sound. Sometimes, a refresh/restart of the arcade player page will bring the sound back into shape, for at least a while.

Why did the Arcade Game start with All Sorts of Weird Graphics?
The JSMESS system provides an as-accurate-as-possible presentation of an arcade machine when it is powered on. A large amount of arcade machines had “boot-up” or “checksum” sequences, where they would show a variety of messages and graphics to indicate the state and quality of the machine. If a ROM chip failed, or a circuit had burned out, various error messages would show and the arcade machine owner or operator would have to do hardware repairs. This situation continues in the emulations, although the machines are generally not going to blow a fuse or lose hardware. That said, there are a very small number of machines that will start up, and then sit at a cryptic operations message, or be awaiting a key. Where possible, the instructions underneath the game’s video window will give information on what key or keys to press to have the game continue to boot up properly.

At the bottom it mentions a Gamepad. Do I need a Gamepad?
Every arcade game can be played using your keyboard; no gamepad or joysticks are needed. That said, it is possible under some circumstances to hook a USB Gamepad to your computer and have it recognized.

]]>
SFLan Information https://help.archive.org/help/sflan-information/ Fri, 15 Mar 2024 20:22:41 +0000 https://help.archive.org/?p=1380 How can I connect to SFLan?
With a laptop: Be in the vicinity of a SFLan node. Associate with it: The SSID is sflanNN, where NN is the number of node, e.g. sflan13. No WEP. You’ll get an IP number assigned via DHCP. With a house: Contact us at info at archive dot org. (Please include your address and a phone number.) Find out if you have line of sight to another SFLan node, buy a node, and we’ll put it on your roof.

What about IP addresses?
SFLan uses real, routable IP addresses. These are usually given out dynically via DHCP. The nodes themselves use static addresses. We can also assign static addresses for servers. For the techies: We use tunneling, layer 2 and layer 3 bridging in parts on the network to make it all appear as a “flat” LAN. There are pros and cons about this approach. It has worked best for us so far. However, it is a moving target, and might change in the future.

I still have more questions, what should I do?
SFLan is a work in progress. If you have more questions, try the SFLan forum. If you still need help, write to info at archive dot org.

I live at 123 Main St at Crossing; do I have line of sight access to a node?
You can try netstumbler or kismet to look for a SFLan ssid.

What is the cost of a node?
The nodes cost $1100, which includes the price of parts and installation. Discounts are potentially available depending on the location.

How can I get a node?
Send an email with your name, exact address and phone number to info at archive dot org. Be sure to write “SFLan node” (or something similar) in the subject line. The information will be passed on to our fantastic installation team who will contact you.

If I get a node, can my neighbors connect also?
Yes, a SFLan node can connect your neighbors and co-condo association members.

What is included in the node?
Most of our nodes are composed of two radios, but some have three. The components are in a weather tight box with a four foot coax cable and two antennas attached. The whole unit is mounted on your roof (generally) on a pole. There is a picture of our lovely 5’3″ spokesmodel holding one here: http://www.archive.org/iathreads/uploaded-files/AstridB-PICT0017.JPG

What are the power requirements of a node?
A node takes on average 5 watts.

What are the connection characteristics of the network?
There are no average characteristics, but 2MBs shared among 20 or so people would be an example.

What is the percentage of uptime?
SFLan is an experimental network, so the uptime varies. Right now uptime averages around 90% or more.

]]>
The Grateful Dead Collection https://help.archive.org/help/the-grateful-dead-collection/ Fri, 15 Mar 2024 20:22:09 +0000 https://help.archive.org/?p=1378 Why are some shows downloadable and others are Stream Only?
Audience-made Grateful Dead concert recordings are available as downloads while available soundboards are accessible in streaming format only. This is done at the request of the band

The Grateful Dead is separated from the Live Music Archive into its own collection (with its own forum) to avoid confusion about lossless availability. The metadata and reviews for shows and recordings, even those not available for regular download, will remain available for those who maintain direct links. No filesets have been deleted from the Archive; certain items are simply not public now. Text files are available at a separate database.

At this time, the Grateful Dead collection is not open to public uploads. The Grateful Dead Internet Archive Project (GDIAP) will continue its direct management of this collection for the time being.

As far as we know, there has been no change to standard GD fan trading. It is common for bands to have policies that differ between fan trading, versus archiving here.

How do I search by date, by year?
On the Collection page you will see the years in the right hand column. Clicking them launches searches for those years.

Where is “on this day in history”?
You can find this in the About tab.

What are some not even available as stream only?
Shows that have been released commercially are not available.

Where is the recording information?
If that information is in an item you will be able to see it on the details page. Recording metadata may include details page you may include: Source – the path from original source to final file format Taped by – the original taper Transferred by – the person(s) who processed the audio to the final file that was uploaded

Why can’t I upload to the Grateful Dead collection?
At this time uploads by the public are prohibited at the request of the band.

]]>