Files, Formats, and Derivatives – Internet Archive Help Center https://help.archive.org How can we help you? Tue, 17 Dec 2024 15:32:45 +0000 en-US hourly 1 https://wordpress.org/?v=6.8 https://help.archive.org/files/2024/03/cropped-Internet-Archive-Logo-White-on-Black-520x520-1-32x32.png Files, Formats, and Derivatives – Internet Archive Help Center https://help.archive.org 32 32 Files, Formats, and Derivatives – file definitions https://help.archive.org/help/files-formats-and-derivatives-file-definitions-2/ Fri, 15 Mar 2024 22:46:32 +0000 https://help.archive.org/?p=896
formatextensionmediatypereview
Metadatafiles.xmlallthe manifest that records all of the files available for this book; also gives 2 checksums and a format definition for each file; provides the only mechanism for validating that the component data has been downloaded successfully
Simple File Verification.sfvallSimple file verification (SFV) is a file format for storing CRC32 checksums of files to verify the integrity of files. SFV is used to verify that a file has not been corrupted, but it does not otherwise verify the file’s authenticity. https://en.wikipedia.org/wiki/Simple_file_verification
Metadatameta.xmlallInternet Archive’s internal “management” metadata; a proprietary XML format, this file includes information about the scan event (date, # of pages, operator, station, etc.), the contributor, basic bib data (title, author, subject, language), and a set of identifiers
Windows Media Audio.wmaaudioWindows Media Audio (WMA) is a series of audio codecs and their corresponding audio coding formats developed by Microsoft. https://en.wikipedia.org/wiki/Windows_Media_Audio
WAVE.wavaudioWaveform Audio File Format (WAV) is an audio file format standard for storing an audio bitstream. https://en.wikipedia.org/wiki/WAV
Ogg Vorbis.oggaudioVorbis is a free and open-source software project. The project produces an audio coding format and software reference encoder/decoder (codec) for lossy audio compression. Vorbis is most commonly used in conjunction with the Ogg container format and it is therefore often referred to as Ogg Vorbis. https://en.wikipedia.org/wiki/Vorbis
VBR MP3.mp3audioVBR (Variable Bitrate) MP3 is a coding format for digital audio https://en.wikipedia.org/wiki/MP3
VBR M3U.m3uaudioVBR (Variable Bitrate) M3U is a computer file format for a multimedia playlist. One common use of the M3U file format is creating a single-entry playlist file pointing to a stream on the Internet. The created file provides easy access to that stream and is often used in downloads from a website, for emailing, and for listening to Internet radio. https://en.wikipedia.org/wiki/M3U
Shorten.shnaudioShorten (SHN) is a file format used for compressing audio data. It is a form of data compression of files and is used to losslessly compress CD-quality audio files. https://en.wikipedia.org/wiki/Shorten_(codec)
MP3.mp3audioMP3 is a coding format for digital audio https://en.wikipedia.org/wiki/MP3
M3U.m3uaudioM3U is a computer file format for a multimedia playlist. One common use of the M3U file format is creating a single-entry playlist file pointing to a stream on the Internet. The created file provides easy access to that stream and is often used in downloads from a website, for emailing, and for listening to Internet radio. https://en.wikipedia.org/wiki/M3U
MP3 Samplesample.mp3audioxLimited length MP3 audio file derived from source audio file. Typically 30 seconds in length.
Flac.flacaudioFLAC is an audio coding format for lossless compression of digital audio, developed by the Xiph.Org Foundation, and is also the name of the free software project producing the FLAC tools, the reference software package that includes a codec implementation. https://en.wikipedia.org/wiki/FLAC
AIFF.aiffaudioAudio Interchange File Format (AIFF) is an audio file format standard used for storing sound data for personal computers and other electronic audio devices. https://en.wikipedia.org/wiki/Audio_Interchange_File_Format
Advanced Audio Coding.m4aaudioAdvanced Audio Coding (AAC) is an audio coding standard for lossy digital audio compression. Designed to be the successor of the MP3 format, AAC generally achieves higher sound quality than MP3 encoders at the same bit rate. https://en.wikipedia.org/wiki/Advanced_Audio_Coding
Spectrogramspectrogram.pngaudioxA visual representation of the spectrum of frequencies of a signal as it varies with time.
Columbia Fingerprint.afpkaudiox“audio fingerprinting” to enable comparing audio tracks together for “the same” tracks or portions of them
Columbia Fingerprintffp.txtaudiox“audio fingerprinting” to enable comparing audio tracks together for “the same” tracks or portions of them
Essentia High GZesshigh.json.gzaudioxhistorical audio format that tried to do analysis like beats-per-minute, deductions of “genre” of music, etc.
Essentia Low GZesslow.json.gzaudioxhistorical audio format that tried to do analysis like beats-per-minute, deductions of “genre” of music, etc.
Flac FingerPrint.ffpaudioxa community-specific checksum for flac files, important to etree community
ZIP.zipdataZIP is an archive file format that supports lossless data compression. A ZIP file may contain one or more files or directories that may have been compressed. https://en.wikipedia.org/wiki/ZIP_(file_format)
Rich Text Format.rtfdataThe Rich Text Format (often abbreviated RTF) is a proprietary document file format. Most word processors are able to read and write some versions of RTF. https://en.wikipedia.org/wiki/Rich_Text_Format
OpenDocument Text Document.odtdataThe Open Document Format for Office Applications (ODF), also known as OpenDocument, is an open standard file format for spreadsheets, charts, presentations and word processing documents using ZIP-compressed XML files https://en.wikipedia.org/wiki/OpenDocument
HTML.htmldataThe HyperText Markup Language or HTML is the standard markup language for documents designed to be displayed in a web browser. It can be assisted by technologies such as Cascading Style Sheets (CSS) and scripting languages such as JavaScript. https://en.wikipedia.org/wiki/HTML
Shockwave.swfdataSWF is an Adobe Flash file format used for multimedia, vector graphics. SWF files can contain animations or applets of varying degrees of interactivity and function. They may also occur in programs, commonly browser games, using ActionScript. https://en.wikipedia.org/wiki/SWF
RAR.rardataRAR is a proprietary archive file format that supports data compression, error recovery and file spanning. https://en.wikipedia.org/wiki/RAR_(file_format)
OpenType Font.otfdataOpenType is a format for scalable computer fonts.. https://en.wikipedia.org/wiki/OpenType
MIDI.middataMIDI is a technical standard that describes a communications protocol, digital interface, and electrical connectors that connect a wide variety of electronic musical instruments, computers, and related audio devices for playing, editing, and recording music. https://en.wikipedia.org/wiki/MIDI
Word Document.docdataMicrosoft Word is a word processing software developed by Microsoft. https://en.wikipedia.org/wiki/Microsoft_Word
Powerpoint.pptdataMicrosoft PowerPoint is a presentation program. PowerPoint was originally designed to provide visuals for group presentations within business organizations, but has come to be very widely used in many other communication situations, both in business and beyond. https://en.wikipedia.org/wiki/Microsoft_PowerPoint
Excel.xlsdataMicrosoft Excel is a spreadsheet developed by Microsoft. https://en.wikipedia.org/wiki/Microsoft_Excel
JSON.jsondataJSON is an open standard file format and data interchange format that uses human-readable text to store and transmit data objects consisting of attribute–value pairs and arrays (or other serializable values). https://en.wikipedia.org/wiki/JSON
TAR.tardataIn computing, tar is a computer software utility for collecting many files into one archive file, often referred to as a tarball, for distribution or backup purposes. https://en.wikipedia.org/wiki/Tar_(computing)
Text.txtdataIn computing, plain text is a loose term for data (e.g. file contents) that represent only characters of readable material but not its graphical representation nor other objects (floating-point numbers, images, etc.). https://en.wikipedia.org/wiki/Plain_text
GZIP.gzdatagzip is a file format and a software application used for file compression and decompression. https://en.wikipedia.org/wiki/Gzip
Flash Video.flvdataFlash Video is a container file format used to deliver digital video content (e.g., TV shows, movies, etc.) over the Internet using Adobe Flash Player version 6 and newer. https://en.wikipedia.org/wiki/Flash_Video
Cascading Style Sheet.cssdataCascading Style Sheets (CSS) is a style sheet language used for describing the presentation of a document written in a markup language such as HTML. https://en.wikipedia.org/wiki/CSS
ISO Image.isodataAn optical disc image (or ISO image, from the ISO 9660 file system used with CD-ROM media) is a disk image that contains everything that would be written to an optical disc, disk sector by disc sector, including the optical disc file system. https://en.wikipedia.org/wiki/Optical_disc_image
Adobe Illustrator.aidataAdobe Illustrator Artwork (AI) is a proprietary file format developed by Adobe Systems for representing single-page vector-based drawings in either the EPS or PDF formats. The .ai filename extension is used by Adobe Illustrator. https://en.wikipedia.org/wiki/Adobe_Illustrator_Artwork
Tab-Separated Values.tsvdataA tab-separated values (TSV) file is a simple text format for storing data in a tabular structure, e.g., a database table or spreadsheet data, and a way of exchanging information between databases. https://en.wikipedia.org/wiki/Tab-separated_values
7Z.7zdata7z is a compressed archive file format that supports several different data compression, encryption and pre-processing algorithms. https://en.wikipedia.org/wiki/7z
Windows Executable.exedata.exe is a common filename extension denoting an executable file (the main execution point of a computer program) for Microsoft Windows. https://en.wikipedia.org/wiki/.exe
Animated GIF.gifimageThe Graphics Interchange Format is a bitmap image format that was developed by a team at the online services provider CompuServe. https://en.wikipedia.org/wiki/GIF
TIFF.tiffimageTag Image File Format, abbreviated TIFF or TIF, is an image file format for storing raster graphics images, popular among graphic artists, the publishing industry, and photographers. https://en.wikipedia.org/wiki/TIFF
PNG.pngimagePortable Network Graphics is a raster-graphics file format that supports lossless data compression. https://en.wikipedia.org/wiki/Portable_Network_Graphics
JPEG.jpgimageJPEG is a commonly used method of lossy compression for digital images, particularly for those images produced by digital photography. https://en.wikipedia.org/wiki/JPEG
JPEG 2000.jp2imageJPEG 2000 (JP2) is an image compression standard and coding system. https://en.wikipedia.org/wiki/JPEG_2000
Web Video Text Tracks.vttmoviesWebVTT (Web Video Text Tracks) is a (W3C standard for displaying timed text in connection with the HTML5 <track> element. https://en.wikipedia.org/wiki/WebVTT
WebM.webmmoviesWebM is an audiovisual media file format. It is primarily intended to offer a royalty-free alternative to use in the HTML5 video and the HTML5 audio elements. https://en.wikipedia.org/wiki/WebM
Ogg Video.ogvmoviesTheora is a free lossy video compression format. It is is most commonly used in conjunction with the Ogg container format. https://en.wikipedia.org/wiki/Theora
Checksums.md5moviesThe MD5 message-digest algorithm is a cryptographically broken but still widely used hash function producing a 128-bit hash value. https://en.wikipedia.org/wiki/MD5
Matroska.mkvmoviesThe Matroska Multimedia Container is a free and open container format, a file format that can hold an unlimited number of video, audio, picture, or subtitle tracks in one file. https://en.wikipedia.org/wiki/Matroska
MPEG4.m4vmoviesThe M4V file format is a video container format developed by Apple and is very similar to the MP4 format. The primary difference is that M4V files may optionally be protected by DRM copy protection. https://en.wikipedia.org/wiki/M4V
QuickTime.movmoviesQuickTime is a video format that is particularly suited for editing, as it is capable of importing and editing in place (without data copying). https://en.wikipedia.org/wiki/QuickTime_File_Format
MPEG4.mpeg4moviesMPEG-4 is a method of defining compression of visual (AV) digital data. https://en.wikipedia.org/wiki/MPEG-4
MPEG2.mpegmoviesMPEG-2 is a standard for “the generic coding of moving pictures and associated audio information”. https://en.wikipedia.org/wiki/MPEG-2
MPEG2.mpgmoviesMPEG-2 is a standard for “the generic coding of moving pictures and associated audio information”. https://en.wikipedia.org/wiki/MPEG-2
512Kb MPEG4512kb.mp4moviesxLow resolution MPEG4 video file
Thumbnailthumb.jpgmoviesxImages of video captured approximated every 30 seconds. They are used in the player scrubber
h.264 IAia.mp4moviesxDerived h.264 file intended to create web-friendly version of uploaded source mp4 that does not meet the minimum criteria for optimal use in the online media player.
Closed Caption Textcc5.txtmoviesxClosed captions text file captured with tv archive recordings
SubRipalign.srtmoviesxClosed Captions in TV Archive items adjusted to better align with the AV
SubRipcc5.srtmoviesxClosed Captions in TV Archive items
Cinepack.avimoviesCinepak is a lossy video codec developed by Peter Barrett at SuperMac Technologies, and released in 1991 with the Video Spigot, and then in 1992 as part of Apple Computer’s QuickTime video suite. https://en.wikipedia.org/wiki/Cinepak
ASRasr.jsmoviesxAutomatic Speech Recognition closed captions. Computer generated from mp3 audio files that are converted to text files.
ASRasr.srtmoviesxAutomatic Speech Recognition closed captions formatted to run in conjucntion with the related video file. Computer generated from mp3 audio files that are converted to text files.
h.264.mp4moviesAdvanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
Windows Media.wmvmoviesAdvanced Systems Format (wmv) is Microsoft’s proprietary digital audio/digital video container format, especially meant for streaming media. https://en.wikipedia.org/wiki/Advanced_Systems_Format
h.264h.264 720Pmoviesx720px1080p h.264 file. Advanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
h.264h.264 HDmoviesx720px1080p h.264 file. Advanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
for tvarchive.xmlmoviesxTV Archive minimal metadata to create full metadata for a show (eg: program title & description, scheduled duration, etc.)
h.264h.264 popcornmoviesxOnline directly in-the-browser user edited audio/video editor files that will playback arbitrary audio & video files, add textual overlays, maps, and more as well
JPEG Thumbthumb.jpgmoviesxA smaller version of various item image files
JSONalign.jsonmoviesxCaptions alignment (audio wave form vs. captions) to reduce the “drift” between what is spoken vs. what got captioned. They can often have 2-10 seconds of distance between displayed words/captions and heard audio
Derivation Rulesrules.confmovies/audioxPrevents lossy derivatives of source data files in audio and video items
Android Package Archive.apksoftwareThe Android Package with the file extension apk is the file format used by the Android operating system, and a number of other Android-based operating systems for distribution and installation of mobile apps, mobile games and middleware. https://en.wikipedia.org/wiki/Apk_(file_format)
Emulator Screenshotscreenshot.pngsoftwarexScreen capture of an emulated computer game
Mac OS X Disk Image.dmgsoftwareApple Disk Image is a disk image format commonly used by the macOS operating system. When opened, an Apple Disk Image is mounted as a volume within the Finder. https://en.wikipedia.org/wiki/Apple_Disk_Image
iOS App Store Package.ipasoftwareAn .ipa (iOS App Store Package) file is an iOS application archive file which stores an iOS app. https://en.wikipedia.org/wiki/.ipa
Amiga Disk File.adfsoftwareAmiga Disk File (ADF) is a file format used by Amiga computers and emulators to store images of floppy disks. https://en.wikipedia.org/wiki/Amiga_Disk_File
Windows Screensaver.scrsoftwareA screensaver is a computer program that blanks the display screen or fills it with moving images or patterns, when the computer has been idle for a designated time. https://en.wikipedia.org/wiki/Screensaver
Log.logtextsxThere are several logs from scanning, republishing, etc. e.g. Cloth Cover Detection Log, various Republisher Logs, and then the plan Log format for Scribe logs.
PDF.pdftextsThe presentation version on BHL in PDF format. Low quality; sufficient for printing and reading text
Metadatareviews.xmltextsxThe meta.xml file contains all of the item-level metadata for reviews
Metadatameta.xmltextsxThe meta.xml file contains all of the item-level metadata for an item (e.g. title, description, creator, etc.).
MARC Binarymarc.xmltextsthe MARC (bibliographic description) data in XML. MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
MARC Binarymeta.mrctextsthe binary MARC record as retrieved using z39.50. MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
Single Page Original JP2 Tarorig_jp2.tartextsSome books are so large that the volume of images exceed the maximum size for a ZIP archive. For these books, the images are compressed and delivered using TAR. These TAR archives average 2.07 gb and occur .39% of the time (738 out of 191,568 books total).
High quality; Best for use and printing of plates, illustrations, detailed figures and tables
DjVu.djvutextsSimilar to PDF, a proprietary compressed document format.
Low quality; sufficient for printing and reading text
Scandatascandata.xmltextsxScandata is an XML file containing specific per-image information, including if the image should be included in any of the produced formats. The module will find, parse and honors these files if they exist.
Text PDF.pdftextsxPortable Document Format files, containing MRC-compressed images and the OCR result as a hidden (selectable, searchable) text layer. (In some cases, the PDF files can have a slightly different suffix, but the extension remains .pdf)
Item Imageitemimage.pngtextsxPNG image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
chOCRchocr.html.gztextsxOCR results with character-level granularity
Dublin Coredc.xmltextsOAI record in Dublin Core (bibliographic description) XML. Dublin Core is a set of metadata elements that provide a small and fundamental group of text elements through which most resources can be described and cataloged; a metadata format for describing resources.
Metadatameta.sqlitetextsxMetadata for file sync via an sqlite database
Name Metadatanames.xmltextslist, by page, of all the scientific names found in the book; presented in xml format
Item Imageitemimage.jpgtextsxJPG image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
Item Tile__ia_thumb.jpgtextsxItem thumbnail image used in search results tiles
Abbyy ZIPabbyy.gztextsGZipped version of the full ABBYY FineReader XML output, which includes all character-level information (confidence, location, etc.)
Item Imageitemimage.giftextsxGIF image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
EPUB.epubtextsEPUB is an e-book file format that uses the “.epub” file extension. The term is short for electronic publication and is sometimes styled ePub. EPUB is supported by many e-readers, and compatible software is available for most smartphones, tablets, and computers.
DAISYtextsDigital accessible information system (DAISY) is a technical standard for digital audiobooks, periodicals, and computerized text. DAISY is designed to be a complete substitute for print material and is specifically designed for use by people with “print disabilities”, including blindness, impaired vision, and dyslexia. https://en.wikipedia.org/wiki/Digital_Accessible_Information_System
Archive BitTorrentarchive.torrenttextsxDerived torrent file that contains files information on files in an item. archive.org does not seed files.
PNGslip.pngtextsBook scanning slips that get uploaded to reserve an identifier so as not to have to wait hours for a full book to upload
hOCRhocr.htmltextsxBarring any failures in the OCR process, after upload, every item will get one or more hocr.html files which represent the results of OCR jobs. Each hocr.html file contains results for all pages in one set of images (book, PDF, or otherwise), with text, bounding boxes, and confidence at the word level.

For those seeking more detailed OCR results, each _hocr.html file should also have a corresponding chocr.html.gz file, with character-level granularity. (The exact meaning of “character” differs, of course, per script or language).
Generic Raw Book Zipimages.ziptextsxA zip imagestack file formatted to derive the files necessary to create a flip book, pdf and other text formats
Single Page Processed JP2 ZIPjp2.ziptextsA ZIP archive of all of the cleaned, cropped, etc. JP2 page images. These are the highest quality, least modified images that are available after the raw/orig file set.
High quality; Best for use and printing of plates, illustrations, detailed figures and tables
Generic Raw Book Tarjp2.tartextsxA tar imagestack file formatted to derive the files necessary to create a flip book, pdf and other text formats
OCR Page Indexhocr_pageindex.json.gztextsxa simple JSON array annotating where each individual page element starts in the hocr.html file, enabling quick fast-forwarding to an individual page without parsing all the XML.
MARC Sourcemetasource.xmltextsa proprietary XML file recording where the MARC record came from (catalog, operator, zquery, etc.) MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
Metadatascandata.xmltextsa proprietary XML file recording information about each page image (handSide, cropBox, original width & height, etc.)
OCR Search Texthocr_searchtext.txt.gztextsxa plaintext file that is ingested by the full text search engine.
Djvu XMLdjvu.xmltextsa modified version of the DjVu XML standard, these files can also be used to read OCR results, but the recommendation is to instead parse the hOCR files.
Page Numbers JSONpage_numbers.jsontextsxA map of page numbers auto-detected in a book. If the confidence score is high enough, they are sometimes added to scandata.xml
JSONevents.jsontextsxA json file containing information about Republisher events. The format is deprecated and no longer used
DjVUTXTdjvu.txttextsa human-readable plaintext version of the generated djvu.xml file. OCR stands for “Optical Character Recognition;” the conversion of images of text into text characters
Biodiversity Heritage Library METSbhlmets.xmltextsxA format created and used by Biodiversity Heritage Library (BHL)
Comic Book RAR.cbrtextsA comic book archive or comic book reader file (also called sequential image file) is a type of archive file for the purpose of sequential viewing of images, commonly for comic books. https://en.wikipedia.org/wiki/Comic_book_archive
Comic Book ZIP.cbztextsA comic book archive or comic book reader file (also called sequential image file) is a type of archive file for the purpose of sequential viewing of images, commonly for comic books. https://en.wikipedia.org/wiki/Comic_book_archive
Grayscale PDFbw.pdftextsA black and white PDF compiled using binarized versions of the images. The binarized images are not made available.
Low; sufficient for printing with low cost of ink or printing only text images
MOBI.mobitexts.mobi is an e-book file format that is primarily used for Kindle e-readers. https://en.wikipedia.org/wiki/Mobipocket
Web ARChive.warcwebThe Web ARChive (WARC) archive format specifies a method for combining multiple digital resources into an aggregate archive file together with related information. https://en.wikipedia.org/wiki/Web_ARChive
Web ARChive GZwarc.gzwebA compressed Web ARChive (WARC) archive using gzip,a file format and a software application used for file compression and decompression. The format specifies a method for combining multiple digital resources into an aggregate archive file together with related information. https://en.wikipedia.org/wiki/Web_ARChive
]]>
Archive BitTorrents https://help.archive.org/help/archive-bittorrents/ Fri, 15 Mar 2024 20:20:20 +0000 https://help.archive.org/?p=1372 How is the Internet Archive using BitTorrent?
As of summer 2012, the Internet Archive is beta-testing the distribution of our public collections via the BitTorrent protocol (as a supplement to traditional HTTP download).

Currently over 1.4 million Archive Items are available via the BitTorrent protocol, comprising almost a petabyte of public domain materials.

As testing continues, more and more content will be made available through Torrents. For the details, see the related FAQ, Details of Archive-made Torrents.

BitTorrent download requires an up-to-date BitTorrent client.

For general information on the BitTorrent protocol, see Wikipedia or BitTorrent.com.

How do I find Torrents on the Archive?
You can search and browse all our Torrents on the Torrents collection homepage (or one of the media-specific subcollections).

To narrow your own Search or Advanced Search query, add format:”Archive BitTorrent” to your search terms, e.g. https://archive.org/search.php?query=’scifi AND mediatype:audio AND format:”Archive BitTorrent”‘.

The most popular and recent Torrents are available on each tracker’s hotlists, e.g. bt1.archive.org Hot List.

How do I edit metadata of my item?
For existing items use this clickpath from the item’s details page:
Edit > change the information > modify/add metadata as desired > click the “Submit” button at the bottom of the page.

You can only modify items that you created.

Can I download only part of an item using an Archive BitTorrent?
Yes, almost all contemporary BitTorrent clients allow you to select which files included in the Torrent are downloaded. And even when you download only one or some files, you get the speed advantages of using the format.

Many show a list of the files contained in the Torrent, and both folders and individual files can be selected or deselected both before, and during, download.

It is recommend, in fact, that you deselect the top-level directory within the Torrent named ._____padding_file if there is one, as this contains unnecessary (empty) Internet Archive padding files.

My Torrent download never completes?
Most likely, you have an out-of-date Torrent for the Item you are trying to download. The first thing to try is re-downloading the Items’ Torrent, and trying again.

Torrents for Items on the Internet Archive can become obsolete when the Item the Torrent is for changes. In that case, some or (more rarely) all of the files within the Torrent will fail to download completely.

This is because our Torrents rely heavily on webseeding (download directly from our servers, when no peers have the files you are seeking). When files on our servers have changed since the Torrent was made, they will not match expected ‘piece hashes’; some BitTorrent clients (e.g. Transmission) will attempt to re-download file pieces from changed files over and over, forever, assuming there was an error in transmission, when in fact the file has changed.

Torrents that never download at all most likely are the result of a different problem, lack of client support for Getright-style webseeding.

My Torrent download never starts?
It’s worth mentioning that some BitTorrent clients take a very long time to begin downloading when relying on webseeding (a common requirement when using Archive BitTorrents). At times downloads can take upwards of several minutes to start.

We’re not sure exactly why; we suspect those clients exhaust all other options, such as DHT, before falling back on webseeds. (We have observed this behavior with Transmission.)

If you download an up-to-date (current) Torrent from the Archive, and it loads into your BitTorrent client, but download never begins, the most likely cause is that you are using a BitTorrent client that does not support Getright-style webseeding.

Our Torrents rely heavily on webseeding (download directly from our servers, when no peers have the files you are seeking). Some BitTorrent clients (e.g. rTorrent) do not support Getright-style webseeding, and will not be able to download un-seeded Internet Archive Torrents.

At the moment, the only solution to this problem is to use a different client.

Another possibility is that your Torrent file is out of date, because the Item has moved to a new server, and your client does not support redirection of our canonical webseeding URL (and no tracked or discoverable peers are seeding the Torrent).

In this case, the problem can be solved by re-downloading the Torrent file.

How do I tell if a Torrent is being seeded?
Current seed and leech counts are displayed for each Archive Torrent on the relevant Item details pages, in parenthesis next to the Torrent link. These values are cached for five minutes or so, and because clients do not always update our trackers regularly, they may be somewhat out of date.

The number of seeders is shown first, and the number of leechers (downloaders without the complete Torrent) second. The seeder number includes ‘webseeds,’ however, which are only usable by BitTorrent clients that support Getright-style webseeding.

Does the Internet Archive run trackers?
Yes, Internet Archive torrents are tracked by bt1.archive.org and bt2.archive.org.

We are using opentracker, which has proven to be highly scalable.

Our trackers are closed (they track our only own torrents).

How do I use Torrents to upload to archive.org?
Retrieval of Torrents is not the best solution for uploading unless you already have an existing mechanism for creating and seeding Torrents.

This capability is not intended as an alternative to our uploader. It merely enables the Archive to capture material already being distributed via BitTorrent.

Torrent retrieval by the Archive works like this:

If a valid .torrent file is uploaded (e.g. through our Uploader) into an item, when that item is derived, we will instantiate a BitTorrent client (Transmission) and attempt to retrieve the Torrent. If the Torrent is successfully retrieved, its contents will be added to the item. ‘Valid’ in this case means, well-formed and seeded.

Our client will attempt to scrape any listed trackers to find seeding peers, but will also attempt to find peers via DHT and can fall back on Getright-style webseeding when possible.

The Torrent file itself is leeched only long enough to retrieve the file; we do not seed the Torrent after retrieval.

However, all items contents, including those retrieved through this method, are made available via the item’s own Archive Torrent. (Because it contains additional contents, this Archive Torrent will, alas, have a different infohash from the original Torrent. So uploading a Torrent to the Archive does not make us a seeder of it.)

Bonus feature: if you have only a magnet link, and not a Torrent file, you can create a dummy .torrent file by pasting that magnet link into a text file and naming it foo.torrent.

If you upload this dummy Torrent file, we’ll detect that you gave us a magnet link and take care of the rest.

Uploading BitTorrents to the Internet Archive

Starting in 2011, the Internet Archive began automatically retrieving BitTorrent files uploaded into most Community collections.

Uploading a Torrent provides a convenient way to upload many files or large contents, provided seeds (including webseeds) are available for the Torrent.

How to prevent an Archive Torrent from being made
Internet Archive BitTorrents are automatically made for community-contributed items in many collections, and automatically updated when item contents or metadata change.

If you prefer that your item not have an Archive Torrent made for it; or that items within a collection you maintain do not, you can prevent Torrents from being made by including the following metadata tag in your item:

noarchivetorrent=true

Note: adding this tag does not remove existing Torrents, those must be removed using the Item Manager item file management tool.

For instructions on how to edit an item or collection’s metadata, see the FAQ Uploading Content.

Why is the Torrent link for an Item lined out (Torrent)?
While an Item is being updated, its Torrent link is temporarily disabled and shown as Torrent.

Changes to an item usually render any existing Archive BitTorrent for it obsolete. Attempts to download obsolete Archive Torrents will usually fail, as described here: My Torrent download never completes?. (Technically, the problem is that when files within an Item change, they can no longer download correctly via webseeding because the piece hashes for updated files change).

The Torrent link will return to normal when the Item finishes updating and the torrent is updated. The Torrent link may be unavailable for a few minutes or a few hours depending on the size of the Item and how busy the Archive processing cluster is (in very rare cases, it might be disabled for a day or more).

Note: obsolete torrents will continue to be tracked by Archive trackers for some time, but will only be retrievable when seeded by peers who have downloaded the referenced version of the item.

What are peers, seeds, leechers, and snatches?
BitTorrent is a peer-to-peer file-sharing protocol facilitated by centralized trackers. The Internet Archive runs several BitTorrent trackers to allow for peer discovery.

Archive trackers track (but do not log or otherwise record) which peers have pieces of which Torrents; real-time statistics are summarized on tracker hotlists for each of our Trackers.

Internet Archive tracker statistics of interest

Peers: the total number of clients known by the tracker to have pieces of a Torrent, i.e. the sum of seeds and leechers.

Seeds: the number of clients known by the tracker to have all of the pieces of a Torrent available, i.e. those which have downloaded the entire Torrent but remain online.

Leechers: the number of clients known by the tracker to have some of the pieces of a Torrent available, i.e. those currently downloading the Torrent.

Snatches: the number of clients known by the tracker to have downloaded a given Torrent, but which are not currently seeding it.

Note: Internet Archive seeder and peer counts include webseeds; these seeds are available only when using clients that support Getright-style webseeding.

My item does not have a torrent. How can I add one?
You would need to send a request to have a torrent generated to info@archive.org. Please include the URL of the item page.

]]>
Files, Formats, and Derivatives – A Basic Guide https://help.archive.org/help/files-formats-and-derivatives-a-basic-guide/ Fri, 15 Mar 2024 19:02:21 +0000 https://help.archive.org/?p=1292 What are all the derived file formats for?

We have three categories for our derive formats:   Data files – for use by various devices for experiencing the media whether it is a book, audio file, movie, etc.    Metadata files – such as XML allow the item page to function on the site.   Research files – such as spectrogram, fingerprint or checksum files.

(Note: The system will not derive a file format that is a duplicate of the source file uploaded. So if an mp3 is uploaded the system will not derive an mp3 from it.)

Tips for formatting files for upload. 

Uploading texts, such as a flip book.

The are two main ways to create a flip book, upload a pdf or using loose images.

1. Uploading a pdf

The easiest way to create a flip book is to upload a well-formatted pdf of hi-res images. 300-600ppi is recommended. Most home printer/scanners do an adequate job. Be sure the dimensions, resolution and compression settings in the pdf are all correct.

2. Making a flip book from images

To make a flip book you need to upload a zip or tar that contains the files. Here’s how:
1. Use only jpg, jpeg, jp2, tif, tiff, png, gif or bmp files. Any combination of them is acceptable.

2. Name your files sequentially. It is best to use the identifier in the name. For example:
000yourfilename.jpg 
001yourfilename.jpg 
002yourfilename.jpg
and so on

3. Create a zip or tar of the files and name that zip file using the page identifier e.g. yourfilename_images.zip or yourfilename_images.tar. If possible use a compression tool that does not add extra extraneous files. Sometimes the compression tool that comes with your computer does this (as with the Mac tool for example.) Terminal or Cygwin work well.

4. Upload the zip or tar file. Be sure to specify a language to help OCR. The system should do the rest.

For more detailed information visit uploading-images-for-text-items/

Movies

It is best to upload the highest resolution file you have. The system can handle most common file formats. We will derive other file formats that are web-friendly such as h.264 mp4. When uploading an mp4 it is recommended to use .mpeg4 as the file name extension. That way the system will create a more web-friendly h.264 mp4 file. Do not upload compressed formats such a .rar as they will not be derived into formats that can be used by the movie player.

Audio

It is best to upload the highest resolution files you have. The system can handle most common file formats. We will derive other file formats such as VBR and MP3.

]]>
Files, Formats, and Derivatives – Tips & Troubleshooting https://help.archive.org/help/files-formats-and-derivatives-tips-troubleshooting/ Fri, 15 Mar 2024 19:01:30 +0000 https://help.archive.org/?p=1290 How do I change the file format specification?

Usually, the system detects the correct format. If it can’t, it may make the wrong choice or specify it as Unknown.

To select a different format, follow these steps starting at the item details page.

For example https://archive.org/details/TestAudioForInternetArchive:

Select Edit

Select Change the Information

Scroll down to the Files, Formats, and Derivatives section. Next, to the file name, there is a dropdown with many formats.

Unless you are certain what the correct format should be, it might be best to send an email regarding your issue to info@archive.org.

My uploading is failing with a message that the file format is bad?

It is possible that the file is corrupt. If possible, you should recreate the file and try to re-upload again.

My item was removed due to malware being detected?

Our virus checker will remove the entire item from the site if a file being uploaded is detected to have malware. If you are adding this to an existing item, you will lose access to the entire item.

Should this happen you may contact us at info@archive.org.

How can I delete unwanted derived files in my items?

The files derived are typically useful for using the materials or are necessary for the item to function on the site.

The only derived files that can be removed and blocked from being created are lossy files such as mp3 in audio items. There is a radio button in the Edit page of items that can be selected to prevent these files from deriving.

Can I search for file information on the site?

File metadata is not indexed in search except for the file formats.

]]>
File Formats https://help.archive.org/help/file-formats/ Fri, 15 Mar 2024 18:41:00 +0000 https://help.archive.org/?p=1247
extensionmediatypedefinition
files.xmlallthe manifest that records all of the files available for this book; also gives 2 checksums and a format definition for each file; provides the only mechanism for validating that the component data has been downloaded successfully
.sfvallSimple file verification (SFV) is a file format for storing CRC32 checksums of files to verify the integrity of files. SFV is used to verify that a file has not been corrupted, but it does not otherwise verify the file’s authenticity. https://en.wikipedia.org/wiki/Simple_file_verification
meta.xmlallInternet Archive’s internal “management” metadata; a proprietary XML format, this file includes information about the scan event (date, # of pages, operator, station, etc.), the contributor, basic bib data (title, author, subject, language), and a set of identifiers
.wmaaudioWindows Media Audio (WMA) is a series of audio codecs and their corresponding audio coding formats developed by Microsoft. https://en.wikipedia.org/wiki/Windows_Media_Audio
.wavaudioWaveform Audio File Format (WAV) is an audio file format standard for storing an audio bitstream. https://en.wikipedia.org/wiki/WAV
.oggaudioVorbis is a free and open-source software project. The project produces an audio coding format and software reference encoder/decoder (codec) for lossy audio compression. Vorbis is most commonly used in conjunction with the Ogg container format and it is therefore often referred to as Ogg Vorbis. https://en.wikipedia.org/wiki/Vorbis
.mp3audioVBR (Variable Bitrate) MP3 is a coding format for digital audio https://en.wikipedia.org/wiki/MP3
.m3uaudioVBR (Variable Bitrate) M3U is a computer file format for a multimedia playlist. One common use of the M3U file format is creating a single-entry playlist file pointing to a stream on the Internet. The created file provides easy access to that stream and is often used in downloads from a website, for emailing, and for listening to Internet radio. https://en.wikipedia.org/wiki/M3U
.shnaudioShorten (SHN) is a file format used for compressing audio data. It is a form of data compression of files and is used to losslessly compress CD-quality audio files. https://en.wikipedia.org/wiki/Shorten_(codec)
.mp3audioMP3 is a coding format for digital audio https://en.wikipedia.org/wiki/MP3
.m3uaudioM3U is a computer file format for a multimedia playlist. One common use of the M3U file format is creating a single-entry playlist file pointing to a stream on the Internet. The created file provides easy access to that stream and is often used in downloads from a website, for emailing, and for listening to Internet radio. https://en.wikipedia.org/wiki/M3U
sample.mp3audioLimited length MP3 audio file derived from source audio file. Typically 30 seconds in length.
.flacaudioFLAC is an audio coding format for lossless compression of digital audio, developed by the Xiph.Org Foundation, and is also the name of the free software project producing the FLAC tools, the reference software package that includes a codec implementation. https://en.wikipedia.org/wiki/FLAC
.aiffaudioAudio Interchange File Format (AIFF) is an audio file format standard used for storing sound data for personal computers and other electronic audio devices. https://en.wikipedia.org/wiki/Audio_Interchange_File_Format
.m4aaudioAdvanced Audio Coding (AAC) is an audio coding standard for lossy digital audio compression. Designed to be the successor of the MP3 format, AAC generally achieves higher sound quality than MP3 encoders at the same bit rate. https://en.wikipedia.org/wiki/Advanced_Audio_Coding
spectrogram.pngaudioA visual representation of the spectrum of frequencies of a signal as it varies with time.
.afpkaudio“audio fingerprinting” to enable comparing audio tracks together for “the same” tracks or portions of them
ffp.txtaudio“audio fingerprinting” to enable comparing audio tracks together for “the same” tracks or portions of them
esshigh.json.gzaudiohistorical audio format that tried to do analysis like beats-per-minute, deductions of “genre” of music, etc.
esslow.json.gzaudiohistorical audio format that tried to do analysis like beats-per-minute, deductions of “genre” of music, etc.
.ffpaudioa community-specific checksum for flac files, important to etree community
.zipdataZIP is an archive file format that supports lossless data compression. A ZIP file may contain one or more files or directories that may have been compressed. https://en.wikipedia.org/wiki/ZIP_(file_format)
.rtfdataThe Rich Text Format (often abbreviated RTF) is a proprietary document file format. Most word processors are able to read and write some versions of RTF. https://en.wikipedia.org/wiki/Rich_Text_Format
.odtdataThe Open Document Format for Office Applications (ODF), also known as OpenDocument, is an open standard file format for spreadsheets, charts, presentations and word processing documents using ZIP-compressed XML files https://en.wikipedia.org/wiki/OpenDocument
.htmldataThe HyperText Markup Language or HTML is the standard markup language for documents designed to be displayed in a web browser. It can be assisted by technologies such as Cascading Style Sheets (CSS) and scripting languages such as JavaScript. https://en.wikipedia.org/wiki/HTML
.swfdataSWF is an Adobe Flash file format used for multimedia, vector graphics. SWF files can contain animations or applets of varying degrees of interactivity and function. They may also occur in programs, commonly browser games, using ActionScript. https://en.wikipedia.org/wiki/SWF
.rardataRAR is a proprietary archive file format that supports data compression, error recovery and file spanning. https://en.wikipedia.org/wiki/RAR_(file_format)
.otfdataOpenType is a format for scalable computer fonts.. https://en.wikipedia.org/wiki/OpenType
.middataMIDI is a technical standard that describes a communications protocol, digital interface, and electrical connectors that connect a wide variety of electronic musical instruments, computers, and related audio devices for playing, editing, and recording music. https://en.wikipedia.org/wiki/MIDI
.docdataMicrosoft Word is a word processing software developed by Microsoft. https://en.wikipedia.org/wiki/Microsoft_Word
.pptdataMicrosoft PowerPoint is a presentation program. PowerPoint was originally designed to provide visuals for group presentations within business organizations, but has come to be very widely used in many other communication situations, both in business and beyond. https://en.wikipedia.org/wiki/Microsoft_PowerPoint
.xlsdataMicrosoft Excel is a spreadsheet developed by Microsoft. https://en.wikipedia.org/wiki/Microsoft_Excel
.jsondataJSON is an open standard file format and data interchange format that uses human-readable text to store and transmit data objects consisting of attribute–value pairs and arrays (or other serializable values). https://en.wikipedia.org/wiki/JSON
.tardataIn computing, tar is a computer software utility for collecting many files into one archive file, often referred to as a tarball, for distribution or backup purposes. https://en.wikipedia.org/wiki/Tar_(computing)
.txtdataIn computing, plain text is a loose term for data (e.g. file contents) that represent only characters of readable material but not its graphical representation nor other objects (floating-point numbers, images, etc.). https://en.wikipedia.org/wiki/Plain_text
.gzdatagzip is a file format and a software application used for file compression and decompression. https://en.wikipedia.org/wiki/Gzip
.flvdataFlash Video is a container file format used to deliver digital video content (e.g., TV shows, movies, etc.) over the Internet using Adobe Flash Player version 6 and newer. https://en.wikipedia.org/wiki/Flash_Video
.cssdataCascading Style Sheets (CSS) is a style sheet language used for describing the presentation of a document written in a markup language such as HTML. https://en.wikipedia.org/wiki/CSS
.isodataAn optical disc image (or ISO image, from the ISO 9660 file system used with CD-ROM media) is a disk image that contains everything that would be written to an optical disc, disk sector by disc sector, including the optical disc file system. https://en.wikipedia.org/wiki/Optical_disc_image
.aidataAdobe Illustrator Artwork (AI) is a proprietary file format developed by Adobe Systems for representing single-page vector-based drawings in either the EPS or PDF formats. The .ai filename extension is used by Adobe Illustrator. https://en.wikipedia.org/wiki/Adobe_Illustrator_Artwork
.tsvdataA tab-separated values (TSV) file is a simple text format for storing data in a tabular structure, e.g., a database table or spreadsheet data, and a way of exchanging information between databases. https://en.wikipedia.org/wiki/Tab-separated_values
.7zdata7z is a compressed archive file format that supports several different data compression, encryption and pre-processing algorithms. https://en.wikipedia.org/wiki/7z
.exedata.exe is a common filename extension denoting an executable file (the main execution point of a computer program) for Microsoft Windows. https://en.wikipedia.org/wiki/.exe
.gifimageThe Graphics Interchange Format is a bitmap image format that was developed by a team at the online services provider CompuServe. https://en.wikipedia.org/wiki/GIF
.tiffimageTag Image File Format, abbreviated TIFF or TIF, is an image file format for storing raster graphics images, popular among graphic artists, the publishing industry, and photographers. https://en.wikipedia.org/wiki/TIFF
.pngimagePortable Network Graphics is a raster-graphics file format that supports lossless data compression. https://en.wikipedia.org/wiki/Portable_Network_Graphics
.jpgimageJPEG is a commonly used method of lossy compression for digital images, particularly for those images produced by digital photography. https://en.wikipedia.org/wiki/JPEG
.jp2imageJPEG 2000 (JP2) is an image compression standard and coding system. https://en.wikipedia.org/wiki/JPEG_2000
.vttmoviesWebVTT (Web Video Text Tracks) is a (W3C standard for displaying timed text in connection with the HTML5 <track> element. https://en.wikipedia.org/wiki/WebVTT
.webmmoviesWebM is an audiovisual media file format. It is primarily intended to offer a royalty-free alternative to use in the HTML5 video and the HTML5 audio elements. https://en.wikipedia.org/wiki/WebM
.ogvmoviesTheora is a free lossy video compression format. It is is most commonly used in conjunction with the Ogg container format. https://en.wikipedia.org/wiki/Theora
.md5moviesThe MD5 message-digest algorithm is a cryptographically broken but still widely used hash function producing a 128-bit hash value. https://en.wikipedia.org/wiki/MD5
.mkvmoviesThe Matroska Multimedia Container is a free and open container format, a file format that can hold an unlimited number of video, audio, picture, or subtitle tracks in one file. https://en.wikipedia.org/wiki/Matroska
.m4vmoviesThe M4V file format is a video container format developed by Apple and is very similar to the MP4 format. The primary difference is that M4V files may optionally be protected by DRM copy protection. https://en.wikipedia.org/wiki/M4V
.movmoviesQuickTime is a video format that is particularly suited for editing, as it is capable of importing and editing in place (without data copying). https://en.wikipedia.org/wiki/QuickTime_File_Format
.mpeg4moviesMPEG-4 is a method of defining compression of visual (AV) digital data. https://en.wikipedia.org/wiki/MPEG-4
.mpegmoviesMPEG-2 is a standard for “the generic coding of moving pictures and associated audio information”. https://en.wikipedia.org/wiki/MPEG-2
.mpgmoviesMPEG-2 is a standard for “the generic coding of moving pictures and associated audio information”. https://en.wikipedia.org/wiki/MPEG-2
512kb.mp4moviesLow resolution MPEG4 video file
thumb.jpgmoviesImages of video captured approximated every 30 seconds. They are used in the player scrubber
ia.mp4moviesDerived h.264 file intended to create web-friendly version of uploaded source mp4 that does not meet the minimum criteria for optimal use in the online media player.
cc5.txtmoviesClosed captions text file captured with tv archive recordings
align.srtmoviesClosed Captions in TV Archive items adjusted to better align with the AV
cc5.srtmoviesClosed Captions in TV Archive items
.avimoviesCinepak is a lossy video codec developed by Peter Barrett at SuperMac Technologies, and released in 1991 with the Video Spigot, and then in 1992 as part of Apple Computer’s QuickTime video suite. https://en.wikipedia.org/wiki/Cinepak
asr.jsmoviesAutomatic Speech Recognition closed captions. Computer generated from mp3 audio files that are converted to text files.
asr.srtmoviesAutomatic Speech Recognition closed captions formatted to run in conjucntion with the related video file. Computer generated from mp3 audio files that are converted to text files.
.mp4moviesAdvanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
.wmvmoviesAdvanced Systems Format (wmv) is Microsoft’s proprietary digital audio/digital video container format, especially meant for streaming media. https://en.wikipedia.org/wiki/Advanced_Systems_Format
h.264 720Pmovies720px1080p h.264 file. Advanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
h.264 HDmovies720px1080p h.264 file. Advanced Video Coding (AVC), also referred to as H.264 or MPEG-4 Part 10, Advanced Video Coding (MPEG-4 AVC), is a video compression standard based on block-oriented, motion-compensated coding. https://en.wikipedia.org/wiki/Advanced_Video_Coding
.xmlmoviesTV Archive minimal metadata to create full metadata for a show (eg: program title & description, scheduled duration, etc.)
h.264 popcornmoviesOnline directly in-the-browser user edited audio/video editor files that will playback arbitrary audio & video files, add textual overlays, maps, and more as well
thumb.jpgmoviesA smaller version of various item image files
align.jsonmoviesCaptions alignment (audio wave form vs. captions) to reduce the “drift” between what is spoken vs. what got captioned. They can often have 2-10 seconds of distance between displayed words/captions and heard audio
rules.confmovies/audioPrevents lossy derivatives of source data files in audio and video items
.apksoftwareThe Android Package with the file extension apk is the file format used by the Android operating system, and a number of other Android-based operating systems for distribution and installation of mobile apps, mobile games and middleware. https://en.wikipedia.org/wiki/Apk_(file_format)
screenshot.pngsoftwareScreen capture of an emulated computer game
.dmgsoftwareApple Disk Image is a disk image format commonly used by the macOS operating system. When opened, an Apple Disk Image is mounted as a volume within the Finder. https://en.wikipedia.org/wiki/Apple_Disk_Image
.ipasoftwareAn .ipa (iOS App Store Package) file is an iOS application archive file which stores an iOS app. https://en.wikipedia.org/wiki/.ipa
.adfsoftwareAmiga Disk File (ADF) is a file format used by Amiga computers and emulators to store images of floppy disks. https://en.wikipedia.org/wiki/Amiga_Disk_File
.scrsoftwareA screensaver is a computer program that blanks the display screen or fills it with moving images or patterns, when the computer has been idle for a designated time. https://en.wikipedia.org/wiki/Screensaver
.logtextsThere are several logs from scanning, republishing, etc. e.g. Cloth Cover Detection Log, various Republisher Logs, and then the plan Log format for Scribe logs.
.pdftextsThe presentation version on BHL in PDF format. Low quality; sufficient for printing and reading text
reviews.xmltextsThe meta.xml file contains all of the item-level metadata for reviews
meta.xmltextsThe meta.xml file contains all of the item-level metadata for an item (e.g. title, description, creator, etc.).
marc.xmltextsthe MARC (bibliographic description) data in XML. MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
meta.mrctextsthe binary MARC record as retrieved using z39.50. MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
orig_jp2.tartextsSome books are so large that the volume of images exceed the maximum size for a ZIP archive. For these books, the images are compressed and delivered using TAR. These TAR archives average 2.07 gb and occur .39% of the time (738 out of 191,568 books total). High quality; Best for use and printing of plates, illustrations, detailed figures and tables
.djvutextsSimilar to PDF, a proprietary compressed document format.Low quality; sufficient for printing and reading text
scandata.xmltextsScandata is an XML file containing specific per-image information, including if the image should be included in any of the produced formats. The module will find, parse and honors these files if they exist.
.pdftextsPortable Document Format files, containing MRC-compressed images and the OCR result as a hidden (selectable, searchable) text layer. (In some cases, the PDF files can have a slightly different suffix, but the extension remains .pdf)
itemimage.pngtextsPNG image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
chocr.html.gztextsOCR results with character-level granularity
dc.xmltextsOAI record in Dublin Core (bibliographic description) XML. Dublin Core is a set of metadata elements that provide a small and fundamental group of text elements through which most resources can be described and cataloged; a metadata format for describing resources.
meta.sqlitetextsMetadata for file sync via an sqlite database
names.xmltextslist, by page, of all the scientific names found in the book; presented in xml format
itemimage.jpgtextsJPG image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
__ia_thumb.jpgtextsItem thumbnail image used in search results tiles
abbyy.gztextsGZipped version of the full ABBYY FineReader XML output, which includes all character-level information (confidence, location, etc.)
itemimage.giftextsGIF image file to be used as the main image in an item page. For audio items it may appear adjacent to the audio player. For collection items it will appear adjacent to the title. It will be used to create the thumbnail image that is used in search results tiles.
.epubtextsEPUB is an e-book file format that uses the “.epub” file extension. The term is short for electronic publication and is sometimes styled ePub. EPUB is supported by many e-readers, and compatible software is available for most smartphones, tablets, and computers.
textsDigital accessible information system (DAISY) is a technical standard for digital audiobooks, periodicals, and computerized text. DAISY is designed to be a complete substitute for print material and is specifically designed for use by people with “print disabilities”, including blindness, impaired vision, and dyslexia. https://en.wikipedia.org/wiki/Digital_Accessible_Information_System
archive.torrenttextsDerived torrent file that contains files information on files in an item. archive.org does not seed files.
slip.pngtextsBook scanning slips that get uploaded to reserve an identifier so as not to have to wait hours for a full book to upload
hocr.htmltextsBarring any failures in the OCR process, after upload, every item will get one or more hocr.html files which represent the results of OCR jobs. Each hocr.html file contains results for all pages in one set of images (book, PDF, or otherwise), with text, bounding boxes, and confidence at the word level. For those seeking more detailed OCR results, each _hocr.html file should also have a corresponding chocr.html.gz file, with character-level granularity. (The exact meaning of “character” differs, of course, per script or language).
images.ziptextsA zip imagestack file formatted to derive the files necessary to create a flip book, pdf and other text formats
jp2.ziptextsA ZIP archive of all of the cleaned, cropped, etc. JP2 page images. These are the highest quality, least modified images that are available after the raw/orig file set. High quality; Best for use and printing of plates, illustrations, detailed figures and tables
jp2.tartextsA tar imagestack file formatted to derive the files necessary to create a flip book, pdf and other text formats
hocr_pageindex.json.gztextsa simple JSON array annotating where each individual page element starts in the hocr.html file, enabling quick fast-forwarding to an individual page without parsing all the XML.
metasource.xmltextsa proprietary XML file recording where the MARC record came from (catalog, operator, zquery, etc.) MARC is a bibliographic data format describing standards for the representation and communication of bibliographic and related information in machine-readable form, and related documentation
scandata.xmltextsa proprietary XML file recording information about each page image (handSide, cropBox, original width & height, etc.)
hocr_searchtext.txt.gztextsa plaintext file that is ingested by the full text search engine.
djvu.xmltextsa modified version of the DjVu XML standard, these files can also be used to read OCR results, but the recommendation is to instead parse the hOCR files.
page_numbers.jsontextsA map of page numbers auto-detected in a book. If the confidence score is high enough, they are sometimes added to scandata.xml
events.jsontextsA json file containing information about Republisher events. The format is deprecated and no longer used
djvu.txttextsa human-readable plaintext version of the generated djvu.xml file. OCR stands for “Optical Character Recognition;” the conversion of images of text into text characters
bhlmets.xmltextsA format created and used by Biodiversity Heritage Library (BHL)
.cbrtextsA comic book archive or comic book reader file (also called sequential image file) is a type of archive file for the purpose of sequential viewing of images, commonly for comic books. https://en.wikipedia.org/wiki/Comic_book_archive
.cbztextsA comic book archive or comic book reader file (also called sequential image file) is a type of archive file for the purpose of sequential viewing of images, commonly for comic books. https://en.wikipedia.org/wiki/Comic_book_archive
bw.pdftextsA black and white PDF compiled using binarized versions of the images. The binarized images are not made available. Low; sufficient for printing with low cost of ink or printing only text images
.mobitexts.mobi is an e-book file format that is primarily used for Kindle e-readers. https://en.wikipedia.org/wiki/Mobipocket
.warcwebThe Web ARChive (WARC) archive format specifies a method for combining multiple digital resources into an aggregate archive file together with related information. https://en.wikipedia.org/wiki/Web_ARChive
warc.gzwebA compressed Web ARChive (WARC) archive using gzip,a file format and a software application used for file compression and decompression. The format specifies a method for combining multiple digital resources into an aggregate archive file together with related information. https://en.wikipedia.org/wiki/Web_AR
]]>