gildas-lormeau/zip.js
 Watch   
 Star   
 Fork   
7 days ago
zip.js

v2.23.0

What's Changed in v2.23.0

New features

  • Add the filter option to the import*() methods of ZipFS and ZipDirectoryEntry: the function receives each entry read from the zip file and returns true (or a promise resolving to true) to import it. It can read the data of the entry to decide, and the entries left out never reach the duplicates policy
  • Add the filter option to the export*() methods and to getExportedSize(): the function receives each ZipEntry of the tree and returns true to export it. Leaving out a directory leaves out its subtree, the entries left out are not read and are not counted by onprogress and onentryprogress
  • Add the filter option to exportFileSystemHandle(), with the same semantics as the zip exports
  • Add the filter option to addFileSystemHandle() and addFileSystemEntry(): the function receives each handle found and its path, so a directory like node_modules or hidden files can be left out without walking them

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.22.1...v2.23.0

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

8 days ago
zip.js

v2.22.1

What's Changed in v2.22.1

Bug fixes

  • A zip file written behind a prefix which was removed together with its whole first entry, like a self-extracting page whose first entry is the face shown by the host, is read again: the entries are shifted to their real positions and their data can be read. Since v2.21.0, the shift was decided on the first central directory record only, and a first record whose local file header lies in the removed bytes proves nothing, so the entries kept their stored offsets and getData() failed with ERR_LOCAL_FILE_HEADER_NOT_FOUND, whatever the strictness option. The reader now checks the following records until one settles the question, and keeps the protection against a damaged end of central directory record whose local file headers are right
  • An entry which lies before the start of the zip file after such a shift is no longer reported as WARNING_UNSORTED_CENTRAL_DIRECTORY

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.22.0...v2.22.1

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

10 days ago
zip.js

v2.22.0

What's Changed in v2.22.0

appendZip()

  • New readerOptions option: the options of the ZipReader which reads the zip file to copy. The zip file used to be read with the default options only, so a zip file the reader rejects by default could not be appended as-is. Set filenameValidation or strictness to copy the entries of a zip file holding unsafe or unusual filenames, filenameEncoding to decode the filenames the duplicate check and the filter option see, and password to let filter read the data of encrypted entries with getData(). The bytes of the entries are copied as-is whatever the options are. A value which is neither an object nor unset throws ERR_INVALID_READER_OPTIONS, now exported by the core builds as well
  • The filter function receives a second argument: the entry of the current zip which has the same filename, as add() or a previous appendZip() call left it, or undefined when there is none. It is the way to apply a duplicate filename policy, since keeping both entries throws ERR_DUPLICATED_NAME: return !existingEntry to keep the entry of the current zip, call remove(existingEntry) and return true to replace it, or compare crc32, uncompressedSize or lastModDate to decide. A removed entry leaves its bytes in the output, as remove() always did, and a strict ZipReader reports them as prepended data. remove() now accepts the EntryMetaData returned by add() in its type declaration
  • The zip file being copied is closed when filter throws, and the documentation of appendZip() now states that add() calls made while filter runs are written before the copied entries, while those made once the copy has started are written after it
  • The entries passed to filter are not modified any more once the callback has returned: the copy used to rewrite their offset with the position in the output and to hang the fields of the rebuilt central directory on them
  • A zip file whose own entries share a filename is rejected with ERR_DUPLICATED_NAME before anything is written, as before, and this is now documented: a ZipWriter holds one entry per filename, so the filter option is the way to keep one of them

Bug fixes

  • new ZipReader(reader, null) reads the zip file with the default options instead of failing with a TypeError when the entries are read; a null options argument is treated as unset, like the other falsy values everywhere in the API

Removed

  • The transferStreams configuration option is removed. It had been a no-op since v2.19.0, when the path transferring the streams to the web workers was removed: the data always crosses the worker boundary chunk by chunk. configure() ignores the key silently, so a call still passing it keeps working; TypeScript reports the key as unknown in WorkerConfiguration, delete it from the call

Tests

  • A real byte overlap between two entries, stretched consistently in the local file header and the central directory record so that the default checkLocalDirectory check passes, is detected with checkOverlappingEntry whatever the order the entries are read in; the existing overlap fixtures overlapped only through a phantom data descriptor

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.21.0...v2.22.0

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

10 days ago
zip.js

v2.21.0

What's Changed in v2.21.0

Bug fixes

  • An entry whose declared data extent (its offset plus its compressed size) ends past the central directory is rejected by getData() with ERR_ENTRY_DATA_OUT_OF_BOUNDS at every strictness level, with or without checkOverlappingEntry. Only the end of the file used to bound it, so a stored entry stretched over the directory returned the directory bytes as its content, and only checkCrc32 could catch it
  • ZipWriter#appendZip() refuses to copy an entry listed with WARNING_MISSING_ZIP64_EXTRA_FIELD, i.e. whose central directory record holds a Zip64 sentinel with no Zip64 extra field resolving it: the method throws ERR_EXTRAFIELD_ZIP64_NOT_FOUND before writing anything, unless the filter option leaves the entry out. Such an entry used to be copied with zero sizes in the rebuilt central directory, or only its local file header through a filter, so the output entry was silently unreadable
  • An entry whose central directory record lacks its Zip64 extra field is now parsed to the end of the record before the defect is reported, so its compression method, dates and other fields are listed with it. An entry whose local file header offset is the unresolved sentinel no longer makes the archive report "prepended data" from the offsets of the other entries
  • The options passed to ZipReader#getEntries() and getEntriesGenerator() now reach the entries: entry.getData() reads them after its own options and before those of the ZipReader constructor, so a strictness, checkLocalDirectory, password or filenameEncoding given to getEntries() applies to the data as the type declarations described
  • A central directory offset stored past the end of the file, e.g. in a damaged or truncated end of central directory record, is reconciled with the directory found before the record instead of failing with ERR_BAD_FORMAT, with the WARNING_MISMATCHED_CENTRAL_DIRECTORY_OFFSET reason deposited and the "strict" level rejecting the archive as before. The same reconciliation now covers a Zip64 end of central directory record stored at the wrong offset: the record is looked for right before its locator, then by its signature within the range the locator allows
  • Before shifting the entries of an archive whose stored central directory offset does not match the directory found, the reader checks that the local file header of the first entry is found at the shifted position and not at the stored one; an archive whose stored offset is short of the directory while its entries sit at the shifted positions is diagnosed as WARNING_PREPENDED_DATA. The shift used to be decided on the direction of the mismatch alone, and a damaged offset could move the entries away from their local file headers
  • An empty archive, i.e. one with no entry, behind prepended data is diagnosed with WARNING_PREPENDED_DATA and its prefix is extracted with extractPrependedData as for a non-empty one; it used to be read without a warning
  • The CRC-32 checksum and the sizes of the data descriptor are compared with the central directory record when the descriptor is read, i.e. when checkOverlappingEntry is set, and a mismatch is reported as WARNING_MISMATCHED_LOCAL_FILE_HEADER_CRC32_OR_SIZES, an error or a warning depending on checkLocalDirectory. A local file header whose CRC-32 checksum and sizes are all zero without the data descriptor flag is tolerated, because some streaming writers leave these fields blank
  • Reading an archive whose end of central directory record points into entry data holding the archive extra data signature no longer takes that data for an encrypted central directory: the record is only looked for at the start of the central directory, where the specification places it
  • EntryMetaData#lastAccessDate and creationDate keep the values of the central directory record when it holds them; the values of the NTFS extra field of the local file header used to overwrite them once the data of the entry was read. When the central directory record holds none, e.g. with the extended timestamp field, whose central form stores the modification time only, the local file header stays their source

appendZip() with the filter option

  • Each kept entry is copied from its local file header up to the exact end of its data or, when it has one, of its data descriptor, whose layout is read back from the zip file. The bytes outside the kept entries are dropped: a self-extracting stub, the data of entries removed earlier and the padding between entries, so a zip file aligned with the usdz option is not aligned any more once filtered. The copy used to stop at the next entry or at the central directory, so the data of a removed entry that followed a kept one was carried into the output
  • Before anything is written, each kept entry is checked to start with a local file header and to end before the next entry or the central directory; otherwise the method throws ERR_LOCAL_FILE_HEADER_NOT_FOUND or ERR_OVERLAPPING_ENTRY and leaves the current zip unchanged. The same checks apply when the output is a split zip file, whose entries are copied one by one as well
  • The general purpose bit flag of a copied entry is written as-is in the rebuilt central directory, including the bits zip.js never sets itself. A filtered copy into a split zip file writer now starts the first disk with the split zip file signature whatever the source starts with
  • A copy that fails after some bytes were written marks the ZipWriter with hasCorruptedEntries and the error with corruptedEntry, as add() does; a failure before the first byte leaves the writer clean. Every filter call completes before any data is copied

Documentation

  • appendZip() states that the comment and the digital signature of the zip file are not copied, since its central directory is rebuilt, and that a source copied as a whole into a split zip file writer carries its self-extracting stub after the split zip file signature of the first disk, where no system runs it; ERR_LOCAL_FILE_HEADER_NOT_FOUND, ERR_OVERLAPPING_ENTRY, ERR_EXTRAFIELD_ZIP64_NOT_FOUND and ERR_ENTRY_DATA_OUT_OF_BOUNDS say when the method and getData() throw them
  • WARNING_MISMATCHED_CENTRAL_DIRECTORY_OFFSET covers both directions of the mismatch and names the local file header check that decides the shift; WARNING_MULTIPLE_END_OF_CENTRAL_DIRECTORY states that it is never deposited as a warning, since "balanced" rejects it like "strict" and "tolerant" reads the last record and reports the stale one as WARNING_TRAILING_CENTRAL_DIRECTORY_DATA; the reasons of ERR_AMBIGUOUS_ARCHIVE list "mismatched central directory offset"
  • checkLocalDirectory describes the data descriptor comparison and the blank local file header tolerance; lastAccessDate and creationDate say which record they are read from

Dependencies

  • The transitive development dependency brace-expansion is updated in the lock file of the benchmarks; no runtime dependency changed, zip.js has none

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.20.0...v2.21.0

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

10 days ago
zip.js

v2.20.0

What's Changed in v2.20.0

New features

  • ZipWriter#appendZip() accepts a filter option: a function called once per entry of the appended zip file, in central directory order, and the entry is copied when it returns or resolves to true. The entries left out leave no bytes behind. The duplicate filename check applies to the kept entries only. Together, this edits an existing zip file into a new one without decompressing its data: appendZip(reader, { filter }) copies the entries to keep as-is, then add() writes the replacements and the additions. With the option, the data is copied entry by entry and the bytes outside the entries, e.g. a self-extracting stub, are not copied. Without it, the zip file is copied as a whole as before. The option also works with split zip file writers

Bug fixes

  • An archive carrying a digital signature record, i.e. written with the signCentralDirectory option, failed to open with ERR_CENTRAL_DIRECTORY_NOT_FOUND once data was prepended to it, e.g. a self-extracting stub: the relocation of the central directory assumed nothing lies between the directory and the end of central directory record. The reader now looks for a signature record ending right before the end of central directory record, and takes the directory to end where that record starts
  • An end of central directory record holding the Zip64 sentinel in its offset, size or disk number field with no Zip64 locator in front of it is rejected with ERR_EOCDR_LOCATOR_ZIP64_NOT_FOUND at every strictness level. It used to be opened with a "prepended data" warning and entries whose data could not be read, since the sentinel was taken for an offset
  • A central directory record holding the Zip64 sentinel in a size or offset field with no Zip64 extra field resolving it no longer makes the whole archive fail getEntries() at the "balanced" and "tolerant" levels: the entry is listed, the new WARNING_MISSING_ZIP64_EXTRA_FIELD reason names it on ZipReader#warnings, reading its data throws ERR_EXTRAFIELD_ZIP64_NOT_FOUND, and the other entries stay readable. The "strict" level keeps throwing ERR_EXTRAFIELD_ZIP64_NOT_FOUND from getEntries(), as every level did before
  • EntryMetaData#lastModDate takes the value of the NTFS extra field (0x000a, 100 ns) over the value of the extended timestamp field (0x5455, 1 s) when a record carries both, as the type declarations already said for rawLastModDate. The extended timestamp used to win because it was read last, so an entry zip.js itself writes with a sub-second date and lastAccessDate, creationDate or ntfsTimestamp: true read back with the milliseconds dropped

Behaviour changes

  • When the offset stored in the end of central directory record points past the central directory actually found, the reader deposits the new WARNING_MISMATCHED_CENTRAL_DIRECTORY_OFFSET reason and the "strict" level rejects the archive with ERR_AMBIGUOUS_ARCHIVE; the relocation used to be silent at every level. This is the shape of an archive written with absolute offsets for a prefix that is no longer there: when the local file header of the first entry is found at the same shifted position, the entries are now read from the shifted positions instead of failing with ERR_LOCAL_FILE_HEADER_NOT_FOUND
  • EntryMetaData#executable is false for directories, as it already was for symbolic links: the execute bits of a directory mean it can be searched, and every directory carries them, so the flag was true for every folder written by zip.js. unixMode still holds the bits

Documentation

  • WARNING_MISMATCHED_CENTRAL_DIRECTORY_OFFSET, WARNING_MISSING_ZIP64_EXTRA_FIELD and ZipWriterAppendZipOptions are documented; ERR_EOCDR_LOCATOR_ZIP64_NOT_FOUND says when it is thrown; rawLastModDate and executable describe the precedence rules above; GetEntriesOptions#filenameValidation states that the filename reported is the decoded central directory name or the one of a valid Unicode Path extra field, and that the name validated is that final name
  • The site documents how to read and write Zstandard entries (compression method 93) with registerCodec(): a codec built on the node:zlib streams for Node.js, the platform CompressionStream classes on Bun, and a JavaScript decoder such as fzstd elsewhere

Benchmarks

  • The ZIP API added to node:zlib in Node.js 26.8 is measured next to jszip, fflate and archiver in BENCHMARKS.md, and the page was re-run on Node.js 26.10, except its browser encryption table, measured on 2026-09-26. The page now says which version of zip.js each table was measured with

Dependencies

  • The transitive development dependencies markdown-it and brace-expansion are updated in the lock file; no runtime dependency changed, zip.js has none

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.19.0...v2.20.0

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

10 days ago
zip.js

v2.19.0

What's Changed in v2.19.0

Bug fixes

  • A Unicode Path extra field (0x7075) no longer bypasses filenameValidation: the name of a valid field replaced the decoded central directory name after that name had been validated, so an entry named safe.txt in its record and ../evil.txt in the field was listed as ../evil.txt at the "strict" and "balanced" levels. The validation and normalizeFilename now run on the final name, so such an entry fails getEntries() with ERR_UNSAFE_FILENAME carrying the name of the field, normalizeFilename receives that name and can repair it, and "tolerant" still keeps it. Conversely, an unsafe record name overridden by a safe field is now accepted, since the name reported is the safe one; rawFilename still holds the bytes of the record
  • With unixExtraFieldType: "unix", the Info-ZIP Unix type 2 extra field (0x7855) is written as Info-ZIP's zip writes it and its extrafld.txt specifies: the 2-byte uid and gid in the local file header, and the tag with a size of 0 in the central directory record, where the 4 bytes used to be repeated. As for the archives Info-ZIP writes, entry.uid and entry.gid of such an entry are undefined after getEntries() and filled in once the entry data has been read: code reading them right after getEntries() on an archive written this way by 2.19.0 has to read the data first, or read entry.localDirectory.uid after getData(), or write the entries with the default "infozip" type. Archives written with "unix" by earlier versions keep reporting the ids at getEntries(), since their central directory copy holds them
  • ZipFS exports keep re-emitting imported ids as a New Unix extra field (0x7875), which stores them in both records, unless unixExtraFieldType is set on the export
  • When a record carries both a New Unix extra field (0x7875) and a type 2 field (0x7855) holding ids, the ids come from the New Unix field, to which Info-ZIP gives precedence and which is not limited to 16 bits; the type 2 ids used to win, so a uid of 70000 next to its truncated 4464 was reported as 4464. extraFieldUnix.uid and extraFieldUnix.gid still report the type 2 values
  • ZipWriter#add() no longer accesses the readable getter of a reader that implements readUint8Array(), e.g. a BlobReader, a Uint8ArrayReader or an HttpRangeReader: the check for a usable reader read readable first, which on these classes creates a new stream on each access, and that stream pulled its first chunk before being discarded, one wasted range request per add() for a remote reader

Behaviour changes

  • The default chunkSize is 256 KiB instead of 64 KiB. Every stage of the pipeline of an entry holds up to one chunk and the data crosses the boundary of a web worker one chunk per message, so the change buys fewer messages for more memory per entry in progress. Where it shows: concurrent add() on Bun 1.4.2 takes 0.24 s instead of 1.20 s, since its native CompressionStream leaves the JavaScript thread only for writes larger than 128 KB, and the worker paths gain in proportion to the number of messages saved. What it costs, measured in isolation on Node.js in-process: about 20 MB more peak memory on a 256 MB stream and on a 20 MB text entry. Applications on memory-tight targets, e.g. browser extensions or mobile pages with many entries in flight, can keep the previous value with configure({ chunkSize: 64 * 1024 }). BENCHMARKS.md was re-run on the new default
  • The streams of an entry are never transferred to a web worker any more: the data always crosses the boundary chunk by chunk, which was never slower and is faster in the browsers, measured with a warmed worker on a 32 MiB deflated entry: Safari 27 reads in 38 ms against 80 with the transfer and writes in 93 against 156, Firefox 156 reads in 75 to 80 ms against 86 to 89 and writes in 32 to 36 against 79 to 87, Chrome 153 writes in 69 against 85 and reads in the same time, Deno is on par, and Bun does not support transferable streams. The transferStreams option of configure() and of the per-call options is accepted and ignored, and deprecated in the type declarations; an application that set transferStreams: false to work around a failure has nothing to change and can drop the option

Documentation

  • GetEntriesOptions#filenameValidation and GetEntriesOptions#normalizeFilename state that the name validated and normalized is the final one, after a valid Unicode Path extra field has replaced it; extraFieldUnix and extraFieldInfoZip of EntryMetaData and of LocalDirectory, and EntryMetaData#localDirectory, describe which field the ids come from and when the local file header fills them in; ZipWriterConstructorOptions#unixExtraFieldType says where the "unix" ids are stored and when a reader sees them; Configuration#chunkSize gains a remark on its memory and message cost; Reader#readable says it returns a new stream reading from the start on each access; WorkerConfiguration#transferStreams is marked deprecated
  • The tests README and two test comments no longer describe streams being transferred to the worker

Tests and continuous integration

  • The filename validation test builds an entry whose Unicode Path extra field escapes the directory and checks it at every level and through normalizeFilename; the Unix extra field layout test pins the empty central directory copy of the type 2 field and reads the ids after the data; the local header ids test adds a record carrying both Unix fields with ids and a central directory carrying only the New Unix field while the local file header carries only the type 2 one; a new test counts the reads a reader receives from add()

Benchmarks

  • QUICK=1 runs the benchmark scripts on a reduced matrix in about 50 seconds, for a smoke check of the harness rather than for the page; bench-crc.js follows the current options of the deflate stream

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.18.2...v2.19.0

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

17 days ago
zip.js

v2.18.2

What's Changed in v2.18.2

Bug fixes

  • With checkOverlappingEntry set, a data descriptor whose CRC-32 disagrees with the central directory, a corrupt CRC-32 for instance, is now reported in localDirectory.dataDescriptor with its own fields, its CRC-32 differing from entry.crc32 as LocalDataDescriptor#crc32 documents. Until now the signed layout was kept only when its CRC-32 and sizes agreed with the central directory, since 2.18.1 among 4- or 8-byte sizes, and otherwise the record was read at the width the Zip64 extra field of either record announces without the signature, i.e. at a layout known not to match, so a signed descriptor with a corrupt CRC-32 came back with signature false, the signature bytes as crc32 and its sizes shifted by four bytes. The layout is now the one whose sizes agree with the central directory, its CRC-32 breaking ties when the central directory stores one; when no layout agrees, the record is read at the announced width as before, with the signature when it starts with one. The tie-break applies to AES entries that store a CRC-32 (AE-1, what WinZip writes for most files) like to any other entry; the CRC-32 of an AES entry used to be ignored. Reading the data of the entry is unchanged, and reading without checkOverlappingEntry never consulted the descriptor

Documentation

  • The remarks of LocalDataDescriptor#signature and LocalDataDescriptor#zip64 describe how the layout is chosen

Tests and continuous integration

  • The data descriptor width test now reads a signed descriptor whose CRC-32 alone is corrupt and one whose compressed size is corrupt, and a new test writes an AE-2 entry with a signed data descriptor, patches it into an AE-1 entry with a disagreeing and then an agreeing descriptor CRC-32, and checks the layout reported for each and for the AE-2 entry

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.18.1...v2.18.2

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

17 days ago
zip.js

v2.18.1

What's Changed in v2.18.1

New features

  • LocalDataDescriptor has a zip64 property, true when the sizes of the data descriptor read with checkOverlappingEntry are stored as 8-byte values

Bug fixes

  • With checkOverlappingEntry set, an entry placed past 4 GB whose data descriptor stores 4-byte sizes now reads; getData() used to fail with ERR_UNSUPPORTED_UINT64. The width of the sizes was taken from the presence of a Zip64 extra field in either record, but neither record tells it reliably: the local file header is written before a streaming writer knows the sizes, and the Zip64 extra field of the central directory record, written last, describes that record, not the descriptor. Go's archive/zip, for instance, gives a small entry placed past 4 GB a Zip64 extra field in its central directory record for the offset alone and a 4-byte descriptor, while its large streamed entries get an 8-byte descriptor with no local Zip64 extra field, which is why the central directory record was consulted; the descriptor of the small entry was therefore read as 8 bytes wide, across the next record. The descriptor is now read with the layout, among the two widths with and without the signature, whose CRC-32, when the entry stores one, and sizes agree with the central directory; when none does, it is read at the width the Zip64 extra field of either record announces, without the signature, as before. Reading without checkOverlappingEntry never depended on the descriptor and is unchanged

Tests and continuous integration

  • A new test builds the archives by hand, behind a reader that fakes a 4 GB prefix so the offsets are real, and reads a 4-byte descriptor below and past 4 GB, signed and unsigned, and an 8-byte descriptor whose Zip64 extra field is in the central directory only or in neither record, each in both read orders, and a descriptor agreeing with no layout

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.18.0...v2.18.1

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

17 days ago
zip.js

v2.18.0

What's Changed in v2.18.0

Bug fixes

  • A Zip64 extra field of a local file header that is too short for the 0xFFFFFFFF sentinels of the header or holds a value above Number.MAX_SAFE_INTEGER no longer fails getData() with ERR_EXTRAFIELD_ZIP64_NOT_FOUND or ERR_UNSUPPORTED_UINT64: the sizes of an entry come from the central directory, so the field is reported as WARNING_MALFORMED_EXTRA_FIELD on entry.warnings and the entry is read. A local file header whose sizes hold the sentinels with no Zip64 extra field behind them, which used to pass silently, is reported the same way. In the three cases an entry without a data descriptor keeps the sentinels as its local sizes, which the local file header check reports as WARNING_MISMATCHED_LOCAL_FILE_HEADER_CRC32_OR_SIZES, i.e. ERR_AMBIGUOUS_ARCHIVE under the default strictness and a warning with strictness: "tolerant". Both errors are still raised by getEntries() for a central directory record whose field is too short or holds such a value, where the sizes and the offset have no other source, and a record whose sizes, offset or disk number hold the sentinel with no Zip64 extra field behind it now fails getEntries() with ERR_EXTRAFIELD_ZIP64_NOT_FOUND too, at every strictness, so none of the entries is listed: it used to be listed with sizes of 4 GB and fail later, with a local file header mismatch or ERR_ENTRY_DATA_OUT_OF_BOUNDS, or with ERR_LOCAL_FILE_HEADER_NOT_FOUND for an offset
  • An AES extra field on a record that is not encrypted and whose compression method is not 99 is ignored and reported as WARNING_MALFORMED_EXTRA_FIELD, on ZipReader#warnings with the filename for the central directory record and on entry.warnings for the local file header, and the entry is read with the method its record declares; the field used to override the method and getData() failed with ERR_UNSUPPORTED_COMPRESSION. An AES extra field shorter than 7 bytes, which was ignored silently, is reported the same way. On an encrypted record the field still overrides the method and the conflict is still rejected with ERR_UNSUPPORTED_COMPRESSION, since the data may be AES behind a wrong method
  • The encrypted flag of an entry follows its central directory record. A local file header whose bit 0 is cleared is still ERR_AMBIGUOUS_ARCHIVE under the default strictness, and a reader with strictness: "tolerant" now decrypts the entry and deposits WARNING_MISMATCHED_LOCAL_FILE_HEADER_BIT_FLAG, where it used to follow the local file header, read the ciphertext as plaintext and fail
  • An entry whose strong encryption bit (bit 6 of the general purpose bit flag) differs between the two records is ERR_AMBIGUOUS_ARCHIVE under the default strictness, like the encrypted bit, and the ERR_UNSUPPORTED_ENCRYPTION check reads the central directory record. A bit set in the local file header only used to reject an encrypted entry, AES or ZipCrypto, as unsupported, and a bit set in the central directory only was ignored; a reader with strictness: "tolerant" now decrypts the first with WARNING_MISMATCHED_LOCAL_FILE_HEADER_BIT_FLAG and reports the second as ERR_UNSUPPORTED_ENCRYPTION
  • The cause of an error raised in a worker keeps its code property, e.g. "Z_MEM_ERROR" on the cause of ERR_CODEC_OUT_OF_MEMORY. A structured clone never copies that property, so it was undefined with workers on every host, while the code of the error itself was already carried by the message posted by the worker

Documentation

  • ERR_EXTRAFIELD_ZIP64_NOT_FOUND, ERR_UNSUPPORTED_UINT64, WARNING_MALFORMED_EXTRA_FIELD and WARNING_MISMATCHED_LOCAL_FILE_HEADER_BIT_FLAG document the cases above, and ERR_INVALID_COMPRESSED_DATA, ERR_INVALID_CRC32, ERR_INVALID_UNCOMPRESSED_SIZE and ERR_CODEC_OUT_OF_MEMORY state more precisely the routes on which they are raised: the ERR_INVALID_CRC32 case of Node.js for bytes trailing the deflate stream with checkCrc32, the gzip route of an AE-2 entry running only on a host without "deflate-raw" whose bundled codec cannot take over either, e.g. because the WebAssembly module failed to load, and, on a host without "deflate-raw", a bundled codec that cannot allocate its state when the entry starts not being reported as ERR_CODEC_OUT_OF_MEMORY, since the native codec takes over through a gzip container unless useCompressionStream is false

Tests and continuous integration

  • New tests cover the three ways a Zip64 extra field of a local file header is unreadable, and their central directory counterparts, the stray AES extra field on a deflated and on an encrypted entry, the encrypted bit and the strong encryption bit disagreeing between the two records, and a worker whose codec cannot allocate its state, a test that runs on the browsers with module workers, Deno and Bun

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.17.0...v2.18.0

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com

19 days ago
zip.js

v2.17.0

What's Changed in v2.17.0

New features

  • ERR_CODEC_OUT_OF_MEMORY is a new exported constant, thrown when the codec of an entry cannot allocate the memory it needs, e.g. when the 16 MB heap of the bundled WebAssembly module is exhausted by too many entries processed concurrently in the same worker, or on the page when workers are off, with the error of the codec as the cause. The failure is recognized by the code property of that error, "Z_MEM_ERROR", which the bundled zlib-streams 1.4.0 sets on its allocation failures and which the native DecompressionStream of Node.js would set. It is raised when reading an entry, both when the codec cannot allocate its state and when it fails while inflating, and when writing an entry whose codec cannot allocate its state; a failure while compressing keeps the error of the codec. Such a failure used to surface as the allocation failed error of the codec, so code matching that message should catch the constant instead

Bug fixes

  • A zip file whose compressedSize overshoots the deflate stream of an entry, which the bundled codecs used to read by dropping the extra bytes, now fails on every codec, as it already did with a native DecompressionStream. The error is ERR_INVALID_COMPRESSED_DATA, with the error of the codec as the cause; on Node.js with checkCrc32 set it is ERR_INVALID_CRC32, since its native inflater reports the extra bytes only once the gzip trailer built by zip.js has been written. The bundled codecs are zlib-streams 1.4.0, the WebAssembly one, and zlib-streams-ts 1.2.0, the pure-JS one of the *-native builds
  • A corrupted deflated entry rejects with ERR_INVALID_COMPRESSED_DATA whatever inflates it, with the error the codec raised as the cause: the TypeError of a native DecompressionStream, whose message depends on the engine, or the error of the bundled WebAssembly or pure-JS codec. Only the message-less error of Node.js used to be mapped, so the same corrupted zip file raised Invalid compressed data on Node.js, corrupt deflate stream on Deno, process error:-3 wherever the WebAssembly codec inflates, e.g. in a browser with useCompressionStream off, and inflate failed on Bun, and a caller comparing against the constant was right on Node.js only. An error raised while the data is being read, by the reader of the zip file or by the decryption of the entry, is not a codec failure and reaches the caller unchanged, with its own cause. On Safari and Bun, whose structured clone drops the cause of an error posted by a worker, the cause is rebuilt from its name and message
  • With checkCrc32 set, a stored uncompressed size larger than the data now fails with ERR_INVALID_CRC32 instead of ERR_INVALID_UNCOMPRESSED_SIZE, on every codec, with the error of the inflater as the cause: the inflater verifies the checksum through a gzip trailer built from the stored CRC-32 and size, and rejects that trailer as a whole. With checkCrc32 off it is still ERR_INVALID_UNCOMPRESSED_SIZE, and a size smaller than the data is ERR_INVALID_UNCOMPRESSED_SIZE either way, as soon as the output exceeds it
  • A signal aborted while an entry is being read by getData() or added by add() rejects the operation with the reason of the signal, whether the compressed data is still being consumed or the content still being written. The signal used to guard the pipe feeding the codec alone, so an abort landing once the input had been consumed, e.g. while a large content was still being written to the writer, was ignored and the operation completed. On the oldest supported engines, which ignore the signal option of pipeTo(), the data is written to the end before the operation is rejected
  • With the passwords and requestPassword options of the filesystem API, the ERR_INVALID_PASSWORD error raised when every candidate has failed, or when requestPassword gives up, carries the error raised by the last candidate as its cause. A ZipCrypto entry whose read fails for a reason other than a false accept, i.e. a wrong password slipping past the one-byte header check of ZipCrypto, as one in 256 does, reports that failure instead of trying the next candidate: only ERR_INVALID_CRC32, ERR_INVALID_COMPRESSED_DATA and ERR_INVALID_UNCOMPRESSED_SIZE count as a wrong password, since such a password produces content that the CRC-32 check or the inflater rejects, while a failure of the reader, e.g. a network error, or an ERR_CODEC_OUT_OF_MEMORY error reaches the caller as-is. Any failure of the read used to count as a wrong password, so a reader failure surfaced as ERR_INVALID_PASSWORD once the candidates were exhausted. A corrupted entry read with the right password is still reported as a wrong password, since nothing tells it from a wrong password passing the check
  • On Chrome before 103 or Node.js before 20.12, whose native inflater lacks "deflate-raw", when the WebAssembly codec cannot take over because its module failed to load, e.g. an *-external build deployed without zip-module.wasm next to it, an entry encrypted with AES that stores no CRC-32 (AE-2, what the writer emits for every encrypted entry) is inflated through a gzip container whose trailer carries the CRC-32 of the output received so far. A valid entry failed with ERR_INVALID_CRC32 now and then, more often on a loaded machine and on Node.js, because the trailer was written on a timing guess. It is now written once the inflater has returned everything, and the entry fails with ERR_INVALID_UNCOMPRESSED_SIZE when it inflates to fewer bytes as well as to more
  • The same wrapper serves the checkCrc32 route on every host, where the rewrite costs about 3 µs per entry with the native codec, measured with the benchmark harness; the tables of BENCHMARKS.md are unchanged
  • The WebAssembly codec (zlib-streams 1.4.0) and the pure-JS codec of the *-native builds (zlib-streams-ts 1.2.0) reject an unknown format with a TypeError when the stream is constructed, as the platform CompressionStream and DecompressionStream do; they used to treat it as "deflate". Nothing changes for a zip file; code that tests a format by constructing the stream now gets the right answer

Documentation

  • ERR_INVALID_COMPRESSED_DATA, ERR_INVALID_CRC32 and ERR_INVALID_UNCOMPRESSED_SIZE document when they are raised, the cause they carry and the routes on which one is raised in place of another, the signal option of the reader and of the writer documents the abort landing during the output, and PasswordCandidatesOptions documents the errors counted as a wrong password and the cause of the final ERR_INVALID_PASSWORD

Tests and continuous integration

  • New tests cover the corrupted deflated entry on every inflate route, with the native and the WebAssembly codecs, with and without workers and with and without checkCrc32, the bytes trailing a deflate stream, the memory failure of the codec on both sides with the heap of the WebAssembly module actually exhausted, the gzip route with a large AE-2 entry read several times, which is what caught the race, the abort landing during the output of a read and of a write, and the reader failure and the cause under the password candidates
  • The password candidates fixture is generated until none of the wrong candidates passes the one-byte check of ZipCrypto by chance, which one in 256 did and failed a job on Chrome 87 once
  • The README of the tests names both compression globals that the polyfill runner removes

Full Changelog: https://github.com/gildas-lormeau/zip.js/compare/v2.16.1...v2.17.0

Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com