Skip to content
Go back

zstd vs lz4 vs gzip vs xz: Measured

By KingPin 18 min read
zstd vs lz4 vs gzip vs xz: Measured
Contents

The short version: zstd -3 is the new default, and xz still owns the high end

We ran 19 compressor configs on 4 data sets, 3 runs each. The headline: zstd -3 beats gzip -6 on every compressible data set. It compresses 4.2x to 6.1x faster, produces smaller files, and decompresses faster too. If you still type tar czf, switch to tar -I 'zstd -3 -T0' -cf and keep your weekend.

The surprise is at the high end. Measured on one thread, xz -6 beats zstd -19 on compress speed. On logs, xz -6 finished in 45 seconds at a ratio of 23.97. zstd -19 needed 211 seconds for 24.05. Same ratio, 4.7x slower. On a database dump and a container filesystem, xz -6 was both smaller and faster. zstd’s real win at high levels is decompression: zstd -19 output unpacks 3.6x to 6.5x faster than xz output.

Two more results worth knowing: bzip2 -9 beats zstd -3 on ratio on every compressible set, though it is the slowest tool here to decompress. And xz -9 costs 675 MiB of RAM to compress for almost nothing over xz -6. Details and commands below. For the everyday zstd how-to (flags, tar, streaming), see Compression in 2026: zstd Changed the Game, which now uses these numbers.

Full example: Clone the working files at github.com/KingPin/sumguy-examples/linux/compression-bench

This is the second post in SumGuy’s Lab: numbers we measured on our own hardware, with the kit published so you can check our work. The box at the top of this page lists the rig and its limits.

The scorecard

Single-thread medians. Ratio is input size divided by output size, so higher is better. Speeds are MiB of input per second.

QuestionWinnerThe numbers
Replace gzip -6zstd -3398 vs 96 MiB/s on logs, ratio 16.16 vs 12.51
Smallest output, any speedxz -9 / xz -6rootfs 4.26 / 4.19 vs zstd -19 at 3.82 (on logs, zstd -19 --long=27 was smallest at 24.23)
Best ratio per second of compress timezstd -9 or xz -6zstd -9 on logs: 20.2 in 3.1 s
Fastest, tolerable ratiolz4 -1478 to 688 MiB/s, ratios 7.96 / 3.33 / 1.95
Fast restores from a max-ratio filezstd -19635 to 1,874 MiB/s vs xz -6 at 103 to 517
Lowest RAM to compressgzipabout 2 MiB
Already-compressed datazstd -3 or lz40.43 s and 0.27 s vs xz -6 at 95 s
Anything bzip2 is still best atnothingslowest to decompress, 20 to 63 MiB/s

How we tested

One machine, everything in Docker, single-thread configs pinned to one CPU, multithreaded configs pinned to four.

The four data sets are 256 MiB each, and all are public, so you can rebuild the exact files. fetch_data.sh downloads them, cuts them to size, and checks them against SHA256SUMS.

SetWhat it isStands in for
logsLoghub “Thunderbird” syslog: sshd, kernel, cron, postfixText logs
dbdumpSimple English Wikipedia SQL table dumps (2026-10-01), MariaDB mysqldump formatA database backup
rootfspostgres:18 image (PG 18.6, amd64) flattened with crane export, pinned by digestContainer images, filesystem backups
videoBig Buck Bunny 1080p H.264 MP4Already-compressed data (control)

Run it yourself:

Terminal window
git clone https://github.com/KingPin/sumguy-examples.git
cd sumguy-examples/linux/compression-bench
./fetch_data.sh # about a minute, a few hundred MB
./bench.sh # about 2 hours, mostly zstd -19 and xz -9
python3 summarize.py results

Wikimedia deletes old dumps after a few months. If the 2026-10-01 file is gone, change the date in fetch_data.sh. The checksum for dbdump will not match, but the shape of the data stays.

The limits, before you quote anything

Result 1: zstd -3 retires gzip -6

This is the one that holds up. Single-thread medians:

Setgzip -6 compresszstd -3 compressgzip -6 ratiozstd -3 ratiogzip -6 decompresszstd -3 decompress
logs96 MiB/s398 MiB/s12.5116.16902 MiB/s1,298 MiB/s
dbdump53 MiB/s228 MiB/s5.115.22410 MiB/s987 MiB/s
rootfs29 MiB/s180 MiB/s2.803.00279 MiB/s726 MiB/s

That is 4.2x, 4.3x, and 6.1x faster compression with a better ratio on all three, and decompression is faster as well. The ratio gain on the database dump is small (5.22 vs 5.11). The speed gain is not.

If your reflex is gzip -9 because “more is better”, skip it. gzip -9 gained 0.5% to 3.8% in ratio over gzip -6 and took 2.4x to 3.4x longer. And zstd -3 is already ahead of it on both axes.

Here is the drop-in change for a tarball backup:

Terminal window
tar -I 'zstd -3 -T0' -cf backup.tar.zst /srv/data
tar -I zstd -xf backup.tar.zst -C /restore

-T0 means “use all physical cores”. The zstd man page defines it that way, and --auto-threads=logical changes it. Plain zstd is single-threaded and defaults to level 3.

Database dumps pipe the same way:

Terminal window
pg_dump mydb | zstd -3 -T0 > mydb.sql.zst
zstd -dc mydb.sql.zst | psql mydb

Result 2: zstd -19 vs xz -6, the upset

Both single-thread, medians:

Setxz -6 timexz -6 ratiozstd -19 timezstd -19 ratio
logs45 s23.97211 s24.05
dbdump83 s7.25135 s6.94
rootfs83 s4.1994 s3.82

On logs the two tie on ratio and xz is 4.7x faster. On the other two sets xz is smaller and faster. At the top of the dial, xz is the better single-thread compressor.

Now the other half. Compression happens once. Decompression happens every time somebody installs the package or restores the archive:

Setxz -6 decompresszstd -19 decompresszstd advantage
logs517 MiB/s1,874 MiB/s3.6x
dbdump171 MiB/s1,108 MiB/s6.5x
rootfs103 MiB/s635 MiB/s6.2x

So the rule is about who reads the file:

Terminal window
# artifact many people download
zstd -19 -T0 --long=27 release.tar -o release.tar.zst
# cold archive nobody will touch for a year
xz -6 -T0 old-logs.tar

The -T0 is not decoration. A bare xz -6 on a current distro already runs multithreaded, and a bare zstd -19 does not, so write the thread count down.

One oddity: the first zstd -19 run on logs took 275 seconds, and runs 2 and 3 took 211 and 202 seconds. That is 30% slower on run 1. We have no explanation, we did not chase one, and the median is what we report.

Result 3: xz -9 is a RAM tax

Setxz -6 ratioxz -9 ratioxz -6 timexz -9 time
logs23.9723.9045 s56 s
dbdump7.257.4183 s123 s
rootfs4.194.2683 s100 s

On logs, -9 was slightly worse than -6. On the other two it gained 2.2% and 1.8%. The memory bill:

On a Raspberry Pi with 1 GB of RAM, that is a bad trade. Nobody will notice a 1.8% smaller backup. Everyone will notice the OOM kill. Use -6. (An xz -9 file also needs those 66 MiB on every machine that unpacks it, which a small box may not have to spare.)

Result 4: bzip2 beat zstd -3, and has no job anyway

On ratio alone, bzip2 -9 beats zstd -3 on all three compressible sets:

Setbzip2 -9 ratiozstd -3 ratio
logs20.5416.16
dbdump6.405.22
rootfs3.123.00

But bzip2 pays for it. It compresses at 9 to 14 MiB/s and decompresses at 20 to 63 MiB/s, the slowest decompression of anything we tested. Compare zstd -9:

No zstd level is beaten by bzip2 on ratio and speed at once. If you have a bzip2 pipeline from 2008, it still works, and replacing it with zstd -9 makes it faster to compress, and zstd -9 or xz -6 makes it much faster to restore. For new work, bzip2 has no job left.

Result 5: lz4 is for when the CPU is the problem

lz4 -1 is fast: 478 to 688 MiB/s on one thread, decompression around 1,470 to 1,700 MiB/s. It is also the worst ratio of anything here:

Setlz4 -1 ratiozstd -1 ratiolz4 -1 compresszstd -1 compress
logs7.9614.74688 MiB/s382 MiB/s
dbdump3.335.17562 MiB/s278 MiB/s
rootfs1.952.63478 MiB/s216 MiB/s

zstd -1 gives up about half the speed and produces files that are 1.3x to 1.9x smaller. That is the right trade for backups, where you pay storage for the life of the file. It is the wrong trade when CPU or latency is the bottleneck. OpenZFS also defaults to lz4 when compression=on (with the lz4_compress feature enabled), and why lz4 is the pick for fast links.

Terminal window
# a filesystem where you want compression to be invisible
zfs set compression=lz4 tank/media
# streaming over a 10 GbE link where gzip would be the bottleneck
tar -cf - /srv/data | lz4 -1 | ssh backup 'lz4 -d | tar -xf - -C /restore'

Result 6: threads

Every tool we tested has a threaded mode. What you get:

ConfigSpeedup, 1 to 4 threads (compressible sets)Ratio change
zstd -33.6x to 4.4x (logs reach 1,739 MiB/s)none
zstd -192.6x to 2.9xnone
xz -62.8x to 3.1xunder 1% worse
pigz -62.7x to 3.7xnegligible
lz4 -12.1x to 2.9xnone

Two details matter:

Bring the cores, but check free -m first.

Result 7: --long=27, small win on small inputs

Long-distance matching lets zstd see repeats that are far apart. On our 256 MiB inputs, zstd -19 --long=27 gained only 0.7% to 1.2% in ratio and added 9% to 18% to compress time (both at 4 threads). The gain would show up on multi-gigabyte VM images or giant tarballs with repeats far apart. A 256 MiB test cannot show that, so we are not claiming it.

The gotcha is on the other end. The zstd decompressor refuses windows bigger than 128 MiB (--long=27) unless you tell it otherwise. We checked: a --long=27 file decompressed fine with a plain zstd -d, while a --long=28 file failed with “Window size larger than maximum” until we added --long=28. Decompressing a --long=27 file used about 132 MiB of RAM.

Terminal window
zstd -19 -T0 --long=27 vm-image.raw -o vm-image.raw.zst
zstd -d --long=27 vm-image.raw.zst -o vm-image.raw

Pass the flag on both sides anyway and keep the two values in sync. A restore that passes --long=27 fails on a --long=30 file, so change both commands together, in the same README.

Result 8: RAM, all in one place

Peak resident memory, compress side:

ConfigMiB
gzipabout 2
xz -04.5
bzip2 -9about 8
lz411 to 19
zstd -338 to 55
xz -695
zstd -19about 215 to 245
xz -9657 to 675
xz -6 -T4477 to 637
zstd -19 -T4560 to 744

If the box is a Pi, a tiny VPS, or an LXC container with a hard memory limit, gzip, zstd -3, and xz -6 all fit. xz -9 and -T4 at high levels are where the OOM killer shows up.

Result 9: the video control

We threw a 1080p H.264 file at every config. Already-compressed data does not compress, and every config landed at 1.00x. bzip2 made it slightly bigger (0.998). The real difference is how long each tool took to find that out:

ConfigTime to learn nothing
lz4 -10.27 s
zstd -30.43 s
gzip -6about 6 s
bzip2 -923 s
zstd -1958 s
xz -695 s
xz -9129 s

At low levels zstd stores incompressible blocks raw and moves on, and lz4 is fast enough that it barely matters. xz and bzip2 grind through every byte. So do not xz your media library, and if a folder is half video and half text, use zstd, which handles the mix without you sorting it:

Terminal window
tar -I 'zstd -3 -T0' -cf mixed.tar.zst ~/media-and-docs

What this means for your homelab

  1. Backups and tarballs: tar -I 'zstd -3 -T0'. It is the new default, and the reason is the first table in this post.

  2. Log rotation: point logrotate at zstd. Config in /etc/logrotate.d/myapp:

    /etc/logrotate.d/myapp
    /var/log/myapp/*.log {
    daily
    rotate 14
    compress
    compresscmd /usr/bin/zstd
    compressoptions -3
    compressext .zst
    uncompresscmd /usr/bin/unzstd
    }

    Logs are the most compressible data in the post (ratio 16.16 at zstd -3), so this is where the switch pays most.

  3. Container images: Docker buildx’s image exporter uses gzip layers by default, and zstd is opt-in:

    Terminal window
    docker buildx build --output type=image,name=registry.example/app:1,compression=zstd,push=true .

    The registry and every client that pulls it must understand zstd layers, so test with a pull first.

  4. Release artifacts and downloads: zstd -19 -T0, because readers decompress 3.6x to 6.5x faster than from xz.

  5. Cold archives: xz -6 -T0. Usually smaller, faster to make, and the slow decompress does not matter if you never read it.

  6. Filesystems and fast links: lz4, as above.

  7. Do not bother with: gzip -9, xz -9 on small hardware, and bzip2 for anything new.

What we did not test

If you rerun the kit on a server or an ARM board and the ordering changes, we want to see it.

Common Questions

Is pigz as fast as zstd?

No. On four threads, pigz -6 compressed at 109 to 256 MiB/s, while zstd -3 on the same four threads reached 643 to 1,739 MiB/s, about 4 to 7 times faster. zstd -3 also produced smaller files on all three compressible sets. Use pigz only when the reader must have gzip format.

Why is xz using all my CPU cores?

Since xz 5.6.0 (2024), multithreaded compression is the default, so a bare xz -6 runs on every core. Older versions used one thread. To get the old behavior, pass -T1 explicitly. To use a fixed number of cores, pass -T4 or any other count.

Does zstd --long need the flag to decompress?

Only above --long=27. The zstd decompressor accepts windows up to 128 MiB by default, so a --long=27 file decompressed with a plain zstd -d in our test. A --long=28 file failed with “Window size larger than maximum” until we passed --long=28 or --memory=256MB. Pass the flag on both sides to be safe.

How much RAM does xz -9 need?

About 675 MiB to compress and 66 MiB to decompress, versus 95 MiB and 10 MiB at xz -6. In our tests, xz -9 gained under 2.5% in ratio over xz -6 on two sets and was slightly worse on the third. Use -6 on any machine with under a few gigabytes of RAM.

Can I compress already compressed files with zstd?

You can, but you will get about 1.00x. A 1080p MP4 compressed to 1.001x with zstd -3, and the run took 0.43 seconds because zstd passes incompressible blocks through quickly. xz spent 95 seconds on the same file for the same result. Skip the step for video and images.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Next Post
Grafana Dashboard Sprawl: 50 Panels to 10

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts