ZIP files explained: compression, encryption and broken archives
9 min read · Updated 4 October 2026
ZIP is the most widely used archive format in the world. Every major operating system can open a ZIP file without extra software, and it is the standard way to send a folder of files by email, download a batch of photos, or distribute software. Many familiar file types are secretly ZIP archives too: Word and Excel documents (.docx, .xlsx), EPUB e-books, Android apps and Java libraries are all ZIP files with a particular internal layout.
Despite this, ZIP files regularly confuse people. Why did my 200 MB folder only shrink to 195 MB? Is a password-protected ZIP safe for confidential documents? Why does this archive say it is "invalid" or "damaged"? This guide answers those questions.
What is inside a ZIP file
A ZIP archive is a container. For each file it stores:
- the file name and folder path, such as
photos/2026/beach.jpg; - the modification date;
- the compressed data;
- a CRC-32 checksum of the original data, so the extractor can detect corruption;
- the compression method used for that file.
At the very end of the archive is the central directory: a table of contents listing every file and where its data starts. Extractors read this table first. That design has a practical consequence covered later: if the end of a ZIP file is missing, the whole archive can appear broken.
Each file inside a ZIP is compressed separately. That makes it quick to extract one file from a large archive, but it means the compressor cannot take advantage of similarities between files. Formats such as .tar.gz and .7z (in its default "solid" mode) compress many files together as one stream, which can be noticeably smaller for large numbers of similar small files, but makes extracting a single file slower.
How compression works
ZIP almost always uses a method called Deflate. Very roughly, it does two things:
- It looks for repeated sequences of bytes. When a sequence has appeared recently, it stores a short reference ("copy 40 bytes from 2,000 bytes back") instead of the bytes themselves.
- It encodes the result with shorter codes for common symbols and longer codes for rare ones.
This is lossless: extracting gives back exactly the original bytes. Nothing is approximated, unlike JPEG or MP3 compression.
How much a file shrinks depends entirely on how much repetition it contains:
| Content | Typical compressed size |
|---|---|
| Plain text, CSV, logs, HTML, source code | 10 to 35% of the original |
| Uncompressed images (BMP, some TIFF) | often 20 to 60% |
| Office documents (DOCX, XLSX) | about 90 to 100% (they are already ZIP files) |
| JPEG, PNG, WebP photos | about 95 to 100% |
| MP3, MP4, MKV, most audio and video | about 98 to 100% |
| Existing ZIP, 7z, RAR archives | about 100% |
Files that are already compressed have had their repetition removed, so Deflate finds almost nothing to work with. Occasionally a compressed version would even be slightly larger; good ZIP tools then store that file with the Store method (no compression) instead.
Worked example
A project folder contains 30 MB of JPEG photos and 10 MB of CSV data, 40 MB in total.
- The JPEGs shrink by perhaps 1 to 3%, to roughly 29.5 MB.
- The CSV files, which are plain text with lots of repeated values, might compress to around 2 MB.
The resulting ZIP is roughly 31.5 MB, a saving of about 21%. Almost all of that saving comes from the CSV files. Re-zipping the archive, or zipping it at a "higher" compression level, will make very little further difference.
If you need to get the photos smaller, compress them as images (lower JPEG quality, smaller dimensions, or a more efficient format), not as files inside a ZIP.
Creating a ZIP that opens everywhere
A few habits avoid problems for the person receiving your archive:
- Put everything in one top-level folder inside the ZIP. Extracting then creates a single folder instead of scattering files into the recipient's Downloads folder.
- Use simple file names. Very long paths, and characters such as
:?*or|, which some systems allow but Windows does not, can cause extraction errors. - Leave out system clutter. Archives made on a Mac can include
__MACOSXfolders and.DS_Storefiles, which appear as junk on other systems. Windows can addThumbs.dbfiles.
The ZIP creator and extractor packages files and folders into a standard ZIP in your browser, without uploading them, and can open ZIP, TAR and GZ archives to extract individual files.
Passwords and encryption
ZIP files can be password-protected, but there are two very different kinds of protection, and the difference matters.
Traditional ZIP encryption ("ZipCrypto"). This is the original scheme from the early 1990s, and it is what many tools still use by default when you "add a password". It is weak. Known attacks can break it quickly, especially when the attacker knows or can guess part of the contents, which is often the case, since many file types begin with predictable headers. Treat ZipCrypto as a lock that keeps out casual snooping, not as real security.
AES encryption. A later extension to the format encrypts each file with AES, usually with a 256-bit key derived from the password. With a strong password, this is genuinely secure. The catch is compatibility: some built-in operating system extractors cannot open AES-encrypted ZIPs, so the recipient may need a dedicated archive program.
Whichever method is used, keep these limits in mind:
- File names are not hidden. Standard ZIP encryption protects the contents of each file but leaves the list of file names, sizes and dates readable. A file called
2026-salary-review-J-Smith.xlsxreveals plenty without being opened. If names are sensitive, zip the files first, then put that ZIP inside a second encrypted archive, or use a format that encrypts names. - The password is the real protection. An AES-encrypted ZIP with the password
summer2026can be guessed. Use a long random password or passphrase. - Send the password separately. Emailing the password in the same message as the archive defeats the purpose. Use a different channel, such as a phone call or a messaging app.
The browser extractor lists encrypted entries and marks them clearly, but it cannot decrypt them; it extracts any unencrypted files in the same archive and leaves the encrypted ones for a desktop archive program.
Size limits and ZIP64
The original ZIP format uses 32-bit fields, which limit it to:
- files of up to 4 GiB each,
- a total archive size of up to 4 GiB,
- 65,535 entries.
The ZIP64 extension removes these limits, and modern tools switch to it automatically when needed. Older software, some embedded devices and a few upload systems do not understand ZIP64, and report such archives as damaged. If you need to send something larger than 4 GiB to someone with old software, split it into several smaller archives.
Why a ZIP file will not open
When an extractor says an archive is "invalid", "corrupt" or "unexpected end of archive", one of these is usually the reason:
1. The download is incomplete. Because the central directory is at the end, a download that stopped at 95% loses the table of contents. Compare the file size with the size shown on the download page, or better, verify a checksum (see below). Download it again.
2. It is not really a ZIP file. A failed download sometimes saves an HTML error page or login page with a .zip name. If the "archive" is only a few kilobytes, open it in a text editor; if you see HTML, the download link needed you to sign in or had expired.
3. It is part of a split archive. Files named archive.z01, archive.z02 … archive.zip, or archive.zip.001, archive.zip.002, are pieces of one archive. You need all the parts in the same folder, and must open the right one (usually the .zip or the .001 part).
4. It uses a compression method your extractor does not support. Besides Deflate, the format allows other methods such as Deflate64, BZIP2, LZMA, Zstandard and others. Archives made by some programs with "maximum" settings may use one of these. A different extractor may succeed, or ask the sender to use standard Deflate.
5. It uses AES encryption and your built-in extractor only supports the older scheme.
6. File names or paths are the problem. Paths longer than Windows' traditional 260-character limit, names containing characters Windows forbids, or names stored in an old character encoding (which appear as garbled symbols) can make extraction fail part-way. Extracting to a short folder path such as C:\x\ often helps with long paths.
7. The data really is corrupted. If the CRC-32 check fails for a file, its data was damaged after the archive was made: on a failing disk, a faulty USB stick or during transfer. Some tools can recover the undamaged files; the damaged one usually needs a fresh copy.
Checking that an archive arrived intact
ZIP's internal CRC-32 checks detect accidental damage to individual files when you extract them. To confirm that the whole archive you received is exactly the one that was sent, before extracting anything, compare a checksum. The sender calculates a SHA-256 hash of the ZIP, the recipient calculates it again, and the two values must match. The checksum verifier calculates SHA-256 and other hashes for files of any size on your own device.
Splitting large archives
If an email service or upload form limits attachments to, say, 25 MB, and your archive is 90 MB, you can split it into parts. The file splitter and joiner cuts any file into pieces of a chosen size and later rejoins them, checking the result with SHA-256. Note that these are plain byte pieces, not ZIP "split volumes": the recipient needs to rejoin the parts with the same kind of tool before opening the ZIP.
Safety: zip bombs and path tricks
Two kinds of malicious archive are worth knowing about:
- A zip bomb is a tiny archive that expands into an enormous amount of data (a few kilobytes that would decompress to gigabytes or more) to fill a disk or crash software. Good extractors check the ratio between compressed and uncompressed size.
- A path traversal archive (sometimes called "zip slip") contains file names such as
../../startup/evil.bat, trying to write files outside the folder you extract into.
The browser extractor blocks both unsafe paths and suspicious compression ratios. More generally, treat archives from unknown senders with the same caution as any other attachment, and do not run programs found inside them unless you trust the source.
Common mistakes
- Zipping photos or videos to "make them smaller", when they barely compress.
- Relying on traditional ZIP passwords for confidential data.
- Sending the password in the same email as the archive.
- Forgetting that file names inside an encrypted ZIP are still visible.
- Opening only one part of a multi-part archive.
FAQ
Does zipping a file reduce its quality? No. ZIP compression is lossless; extracted files are byte-for-byte identical to the originals.
Why is my ZIP bigger than the original file? With already-compressed data, the archive adds a little overhead for file names and headers, so a ZIP of one JPEG can be slightly larger than the JPEG itself.
What is the difference between ZIP and .tar.gz? TAR bundles files into one stream without compressing them; gzip then compresses that whole stream. The result is often smaller for many small text files, and it preserves Unix file permissions, but you cannot extract one file without reading through the archive. ZIP compresses each file separately and is more widely supported on Windows and macOS.