The short version: zstd -3 is the new default, and xz still owns the high end
We ran 19 compressor configs on 4 data sets, 3 runs each. The headline: zstd -3 beats gzip -6 on every compressible data set. It compresses 4.2x to 6.1x faster, produces smaller files, and decompresses faster too. If you still type tar czf, switch to tar -I 'zstd -3 -T0' -cf and keep your weekend.
The surprise is at the high end. Measured on one thread, xz -6 beats zstd -19 on compress speed. On logs, xz -6 finished in 45 seconds at a ratio of 23.97. zstd -19 needed 211 seconds for 24.05. Same ratio, 4.7x slower. On a database dump and a container filesystem, xz -6 was both smaller and faster. zstd’s real win at high levels is decompression: zstd -19 output unpacks 3.6x to 6.5x faster than xz output.
Two more results worth knowing: bzip2 -9 beats zstd -3 on ratio on every compressible set, though it is the slowest tool here to decompress. And xz -9 costs 675 MiB of RAM to compress for almost nothing over xz -6. Details and commands below. For the everyday zstd how-to (flags, tar, streaming), see Compression in 2026: zstd Changed the Game, which now uses these numbers.
Full example: Clone the working files at github.com/KingPin/sumguy-examples/linux/compression-bench
This is the second post in SumGuy’s Lab: numbers we measured on our own hardware, with the kit published so you can check our work. The box at the top of this page lists the rig and its limits.
The scorecard
Single-thread medians. Ratio is input size divided by output size, so higher is better. Speeds are MiB of input per second.
| Question | Winner | The numbers |
|---|---|---|
| Replace gzip -6 | zstd -3 | 398 vs 96 MiB/s on logs, ratio 16.16 vs 12.51 |
| Smallest output, any speed | xz -9 / xz -6 | rootfs 4.26 / 4.19 vs zstd -19 at 3.82 (on logs, zstd -19 --long=27 was smallest at 24.23) |
| Best ratio per second of compress time | zstd -9 or xz -6 | zstd -9 on logs: 20.2 in 3.1 s |
| Fastest, tolerable ratio | lz4 -1 | 478 to 688 MiB/s, ratios 7.96 / 3.33 / 1.95 |
| Fast restores from a max-ratio file | zstd -19 | 635 to 1,874 MiB/s vs xz -6 at 103 to 517 |
| Lowest RAM to compress | gzip | about 2 MiB |
| Already-compressed data | zstd -3 or lz4 | 0.43 s and 0.27 s vs xz -6 at 95 s |
| Anything bzip2 is still best at | nothing | slowest to decompress, 20 to 63 MiB/s |
How we tested
One machine, everything in Docker, single-thread configs pinned to one CPU, multithreaded configs pinned to four.
- Hardware: Intel Core i7-11800H laptop (8 cores, 16 threads), 64 GB RAM, Ubuntu 22.04, kernel 6.8, Docker 29.8, CPU governor
powersave(intel_pstate). - Pinning: single-thread configs run on one CPU. Four-thread configs run on 4 CPUs on 4 separate physical cores (CPUs 4 to 7 on this chip, where CPU n and n+8 share a core), with the hyperthread siblings left idle.
- Storage: inputs and outputs live in tmpfs (RAM), and decompression writes to
/dev/null. No disk speed in any number. - Method: 3 runs per config, median reported. The first run of every config is decompressed again and compared byte for byte with the input.
- Tools: zstd 1.5.7, lz4 1.10.0, xz 5.8.4, gzip 1.14, pigz 2.8, bzip2 1.0.8. These are the Debian forky packages, and all were the current upstream releases on 2026-10-06, the day we tested.
- Explicit threads: the kit always passes a thread count. Since xz 5.6.0 (2024), a bare
xz -6uses every core by default, and lz4 1.10.0 does the same. If you benchmark with defaults, you are comparing one tool on one core against another on sixteen.
The four data sets are 256 MiB each, and all are public, so you can rebuild the exact files. fetch_data.sh downloads them, cuts them to size, and checks them against SHA256SUMS.
| Set | What it is | Stands in for |
|---|---|---|
| logs | Loghub “Thunderbird” syslog: sshd, kernel, cron, postfix | Text logs |
| dbdump | Simple English Wikipedia SQL table dumps (2026-10-01), MariaDB mysqldump format | A database backup |
| rootfs | postgres:18 image (PG 18.6, amd64) flattened with crane export, pinned by digest | Container images, filesystem backups |
| video | Big Buck Bunny 1080p H.264 MP4 | Already-compressed data (control) |
Run it yourself:
git clone https://github.com/KingPin/sumguy-examples.gitcd sumguy-examples/linux/compression-bench./fetch_data.sh # about a minute, a few hundred MB./bench.sh # about 2 hours, mostly zstd -19 and xz -9python3 summarize.py resultsWikimedia deletes old dumps after a few months. If the 2026-10-01 file is gone, change the date in fetch_data.sh. The checksum for dbdump will not match, but the shape of the data stays.
The limits, before you quote anything
- One machine, one CPU model. A laptop in
powersavegovernor is not a server. Absolute MiB/s will move on your hardware. The ratios between configs travel better. - The logs set is a 2005 supercomputer’s syslog. It is very repetitive, so its ratios run higher than a typical homelab’s mixed app logs.
- 256 MiB inputs. That is too small to show what long-range matching buys on multi-gigabyte files.
- Decompression of the fast tools takes 0.08 to 0.35 seconds, and that includes process start-up. Treat those numbers as a floor, and treat differences under about 20% between lz4 and zstd decompression as noise.
- RAM-backed storage removes the disk. On a real backup job the disk or the network may cap you long before zstd -3 does.
- We tested what is in the table. Other zstd levels,
--fast, zstd dictionaries,xzpresets with-e, and brotli were not tested.
Result 1: zstd -3 retires gzip -6
This is the one that holds up. Single-thread medians:
| Set | gzip -6 compress | zstd -3 compress | gzip -6 ratio | zstd -3 ratio | gzip -6 decompress | zstd -3 decompress |
|---|---|---|---|---|---|---|
| logs | 96 MiB/s | 398 MiB/s | 12.51 | 16.16 | 902 MiB/s | 1,298 MiB/s |
| dbdump | 53 MiB/s | 228 MiB/s | 5.11 | 5.22 | 410 MiB/s | 987 MiB/s |
| rootfs | 29 MiB/s | 180 MiB/s | 2.80 | 3.00 | 279 MiB/s | 726 MiB/s |
That is 4.2x, 4.3x, and 6.1x faster compression with a better ratio on all three, and decompression is faster as well. The ratio gain on the database dump is small (5.22 vs 5.11). The speed gain is not.
If your reflex is gzip -9 because “more is better”, skip it. gzip -9 gained 0.5% to 3.8% in ratio over gzip -6 and took 2.4x to 3.4x longer. And zstd -3 is already ahead of it on both axes.
Here is the drop-in change for a tarball backup:
tar -I 'zstd -3 -T0' -cf backup.tar.zst /srv/datatar -I zstd -xf backup.tar.zst -C /restore-T0 means “use all physical cores”. The zstd man page defines it that way, and --auto-threads=logical changes it. Plain zstd is single-threaded and defaults to level 3.
Database dumps pipe the same way:
pg_dump mydb | zstd -3 -T0 > mydb.sql.zstzstd -dc mydb.sql.zst | psql mydbResult 2: zstd -19 vs xz -6, the upset
Both single-thread, medians:
| Set | xz -6 time | xz -6 ratio | zstd -19 time | zstd -19 ratio |
|---|---|---|---|---|
| logs | 45 s | 23.97 | 211 s | 24.05 |
| dbdump | 83 s | 7.25 | 135 s | 6.94 |
| rootfs | 83 s | 4.19 | 94 s | 3.82 |
On logs the two tie on ratio and xz is 4.7x faster. On the other two sets xz is smaller and faster. At the top of the dial, xz is the better single-thread compressor.
Now the other half. Compression happens once. Decompression happens every time somebody installs the package or restores the archive:
| Set | xz -6 decompress | zstd -19 decompress | zstd advantage |
|---|---|---|---|
| logs | 517 MiB/s | 1,874 MiB/s | 3.6x |
| dbdump | 171 MiB/s | 1,108 MiB/s | 6.5x |
| rootfs | 103 MiB/s | 635 MiB/s | 6.2x |
So the rule is about who reads the file:
- Compress once, read often (release artifacts, package repos, container layers pulled by many hosts):
zstd -19is worth the compress time, because every reader gets a 3.6x to 6.5x faster unpack. - Archive once, read rarely (cold backups, old logs):
xz -6makes smaller files in less time.
# artifact many people downloadzstd -19 -T0 --long=27 release.tar -o release.tar.zst
# cold archive nobody will touch for a yearxz -6 -T0 old-logs.tarThe -T0 is not decoration. A bare xz -6 on a current distro already runs multithreaded, and a bare zstd -19 does not, so write the thread count down.
One oddity: the first zstd -19 run on logs took 275 seconds, and runs 2 and 3 took 211 and 202 seconds. That is 30% slower on run 1. We have no explanation, we did not chase one, and the median is what we report.
Result 3: xz -9 is a RAM tax
| Set | xz -6 ratio | xz -9 ratio | xz -6 time | xz -9 time |
|---|---|---|---|---|
| logs | 23.97 | 23.90 | 45 s | 56 s |
| dbdump | 7.25 | 7.41 | 83 s | 123 s |
| rootfs | 4.19 | 4.26 | 83 s | 100 s |
On logs, -9 was slightly worse than -6. On the other two it gained 2.2% and 1.8%. The memory bill:
- Compress: 675 MiB at
-9vs 95 MiB at-6. - Decompress: 66 MiB at
-9vs 10 MiB at-6.
On a Raspberry Pi with 1 GB of RAM, that is a bad trade. Nobody will notice a 1.8% smaller backup. Everyone will notice the OOM kill. Use -6. (An xz -9 file also needs those 66 MiB on every machine that unpacks it, which a small box may not have to spare.)
Result 4: bzip2 beat zstd -3, and has no job anyway
On ratio alone, bzip2 -9 beats zstd -3 on all three compressible sets:
| Set | bzip2 -9 ratio | zstd -3 ratio |
|---|---|---|
| logs | 20.54 | 16.16 |
| dbdump | 6.40 | 5.22 |
| rootfs | 3.12 | 3.00 |
But bzip2 pays for it. It compresses at 9 to 14 MiB/s and decompresses at 20 to 63 MiB/s, the slowest decompression of anything we tested. Compare zstd -9:
- 3.7x to 9.4x faster to compress than bzip2.
- Within 2% of bzip2’s ratio on logs (20.21 vs 20.54), and ahead of it on rootfs (3.34 vs 3.12).
- Behind on dbdump (5.88 vs 6.40). That is the only set where bzip2 keeps a real ratio lead over zstd -9.
No zstd level is beaten by bzip2 on ratio and speed at once. If you have a bzip2 pipeline from 2008, it still works, and replacing it with zstd -9 makes it faster to compress, and zstd -9 or xz -6 makes it much faster to restore. For new work, bzip2 has no job left.
Result 5: lz4 is for when the CPU is the problem
lz4 -1 is fast: 478 to 688 MiB/s on one thread, decompression around 1,470 to 1,700 MiB/s. It is also the worst ratio of anything here:
| Set | lz4 -1 ratio | zstd -1 ratio | lz4 -1 compress | zstd -1 compress |
|---|---|---|---|---|
| logs | 7.96 | 14.74 | 688 MiB/s | 382 MiB/s |
| dbdump | 3.33 | 5.17 | 562 MiB/s | 278 MiB/s |
| rootfs | 1.95 | 2.63 | 478 MiB/s | 216 MiB/s |
zstd -1 gives up about half the speed and produces files that are 1.3x to 1.9x smaller. That is the right trade for backups, where you pay storage for the life of the file. It is the wrong trade when CPU or latency is the bottleneck. OpenZFS also defaults to lz4 when compression=on (with the lz4_compress feature enabled), and why lz4 is the pick for fast links.
# a filesystem where you want compression to be invisiblezfs set compression=lz4 tank/media
# streaming over a 10 GbE link where gzip would be the bottlenecktar -cf - /srv/data | lz4 -1 | ssh backup 'lz4 -d | tar -xf - -C /restore'Result 6: threads
Every tool we tested has a threaded mode. What you get:
| Config | Speedup, 1 to 4 threads (compressible sets) | Ratio change |
|---|---|---|
| zstd -3 | 3.6x to 4.4x (logs reach 1,739 MiB/s) | none |
| zstd -19 | 2.6x to 2.9x | none |
| xz -6 | 2.8x to 3.1x | under 1% worse |
| pigz -6 | 2.7x to 3.7x | negligible |
| lz4 -1 | 2.1x to 2.9x | none |
Two details matter:
- The
xz -T4output also decompresses faster: 1.9x to 3.3x, because multithreaded xz writes independent blocks. Single-threadxz -6writes one block, and per the xz docs a single-block file cannot be decompressed in parallel. - The memory goes up.
zstd -19 -T4used 560 to 744 MiB to compress, andxz -6 -T4used 477 to 637 MiB. Four threads on a 1 GB box is not free.
Bring the cores, but check free -m first.
Result 7: --long=27, small win on small inputs
Long-distance matching lets zstd see repeats that are far apart. On our 256 MiB inputs, zstd -19 --long=27 gained only 0.7% to 1.2% in ratio and added 9% to 18% to compress time (both at 4 threads). The gain would show up on multi-gigabyte VM images or giant tarballs with repeats far apart. A 256 MiB test cannot show that, so we are not claiming it.
The gotcha is on the other end. The zstd decompressor refuses windows bigger than 128 MiB (--long=27) unless you tell it otherwise. We checked: a --long=27 file decompressed fine with a plain zstd -d, while a --long=28 file failed with “Window size larger than maximum” until we added --long=28. Decompressing a --long=27 file used about 132 MiB of RAM.
zstd -19 -T0 --long=27 vm-image.raw -o vm-image.raw.zstzstd -d --long=27 vm-image.raw.zst -o vm-image.rawPass the flag on both sides anyway and keep the two values in sync. A restore that passes --long=27 fails on a --long=30 file, so change both commands together, in the same README.
Result 8: RAM, all in one place
Peak resident memory, compress side:
| Config | MiB |
|---|---|
| gzip | about 2 |
| xz -0 | 4.5 |
| bzip2 -9 | about 8 |
| lz4 | 11 to 19 |
| zstd -3 | 38 to 55 |
| xz -6 | 95 |
| zstd -19 | about 215 to 245 |
| xz -9 | 657 to 675 |
| xz -6 -T4 | 477 to 637 |
| zstd -19 -T4 | 560 to 744 |
If the box is a Pi, a tiny VPS, or an LXC container with a hard memory limit, gzip, zstd -3, and xz -6 all fit. xz -9 and -T4 at high levels are where the OOM killer shows up.
Result 9: the video control
We threw a 1080p H.264 file at every config. Already-compressed data does not compress, and every config landed at 1.00x. bzip2 made it slightly bigger (0.998). The real difference is how long each tool took to find that out:
| Config | Time to learn nothing |
|---|---|
| lz4 -1 | 0.27 s |
| zstd -3 | 0.43 s |
| gzip -6 | about 6 s |
| bzip2 -9 | 23 s |
| zstd -19 | 58 s |
| xz -6 | 95 s |
| xz -9 | 129 s |
At low levels zstd stores incompressible blocks raw and moves on, and lz4 is fast enough that it barely matters. xz and bzip2 grind through every byte. So do not xz your media library, and if a folder is half video and half text, use zstd, which handles the mix without you sorting it:
tar -I 'zstd -3 -T0' -cf mixed.tar.zst ~/media-and-docsWhat this means for your homelab
-
Backups and tarballs:
tar -I 'zstd -3 -T0'. It is the new default, and the reason is the first table in this post. -
Log rotation: point logrotate at zstd. Config in
/etc/logrotate.d/myapp:/etc/logrotate.d/myapp /var/log/myapp/*.log {dailyrotate 14compresscompresscmd /usr/bin/zstdcompressoptions -3compressext .zstuncompresscmd /usr/bin/unzstd}Logs are the most compressible data in the post (ratio 16.16 at
zstd -3), so this is where the switch pays most. -
Container images: Docker buildx’s image exporter uses gzip layers by default, and zstd is opt-in:
Terminal window docker buildx build --output type=image,name=registry.example/app:1,compression=zstd,push=true .The registry and every client that pulls it must understand zstd layers, so test with a pull first.
-
Release artifacts and downloads:
zstd -19 -T0, because readers decompress 3.6x to 6.5x faster than from xz. -
Cold archives:
xz -6 -T0. Usually smaller, faster to make, and the slow decompress does not matter if you never read it. -
Filesystems and fast links: lz4, as above.
-
Do not bother with:
gzip -9,xz -9on small hardware, and bzip2 for anything new.
What we did not test
- Other zstd levels beyond 1, 3, 9, and 19, and
--fastlevels. - Dictionaries, which can help a lot with many small files.
- brotli, lzip, lzma-alone, and other compressors.
- Real disks or networks. Everything lived in RAM.
- Inputs bigger than 256 MiB, where
--longand block sizes may change the picture. - A different CPU. AVX2, cache size, and the
powersavegovernor all move absolute speeds.
If you rerun the kit on a server or an ARM board and the ordering changes, we want to see it.
Common Questions
Is pigz as fast as zstd?
No. On four threads, pigz -6 compressed at 109 to 256 MiB/s, while zstd -3 on the same four threads reached 643 to 1,739 MiB/s, about 4 to 7 times faster. zstd -3 also produced smaller files on all three compressible sets. Use pigz only when the reader must have gzip format.
Why is xz using all my CPU cores?
Since xz 5.6.0 (2024), multithreaded compression is the default, so a bare xz -6 runs on every core. Older versions used one thread. To get the old behavior, pass -T1 explicitly. To use a fixed number of cores, pass -T4 or any other count.
Does zstd --long need the flag to decompress?
Only above --long=27. The zstd decompressor accepts windows up to 128 MiB by default, so a --long=27 file decompressed with a plain zstd -d in our test. A --long=28 file failed with “Window size larger than maximum” until we passed --long=28 or --memory=256MB. Pass the flag on both sides to be safe.
How much RAM does xz -9 need?
About 675 MiB to compress and 66 MiB to decompress, versus 95 MiB and 10 MiB at xz -6. In our tests, xz -9 gained under 2.5% in ratio over xz -6 on two sets and was slightly worse on the third. Use -6 on any machine with under a few gigabytes of RAM.
Can I compress already compressed files with zstd?
You can, but you will get about 1.00x. A 1080p MP4 compressed to 1.001x with zstd -3, and the run took 0.43 seconds because zstd passes incompressible blocks through quickly. xz spent 95 seconds on the same file for the same result. Skip the step for video and images.