The short version: nginx is faster, and that is not the story
We pinned nginx and Caddy to the same two CPU cores and hit them with the same load. nginx is faster per core on every keep-alive file test, by 1.2x to 5.1x. At homelab traffic, that gap does not matter. Caddy on 2 cores still served 21,000 to 29,000 small files per second, and your Jellyfin box does not see that on its best day.
The results that matter are about connections to the backend, and both servers ship a default worth changing. The common nginx form, proxy_pass pointed straight at a backend hostname, opens a new TCP connection for every request. Under sustained load it runs out of local ports and answers with 502 Bad Gateway. On nginx 1.31.6, the first 28,231 requests worked and everything after that failed. Caddy reuses backend connections, but it keeps only 32 idle ones per backend. Past 32 concurrent requests it slowly leaks ports too. Over 90 seconds at 100 concurrent requests, its TIME_WAIT sockets filled the port range and its throughput dropped to 6,885 requests per second. Each fix is a few lines, shown below.
Full example: Clone the working files at github.com/KingPin/sumguy-examples/networking/nginx-vs-caddy-bench
This is the first post in SumGuy’s Lab: posts built on numbers we measured on our own hardware, with the test kit published so you can check our work. The box at the top of this page lists the rig and its limits.
The scorecard
| Test (2 cores) | nginx 1.31.6 | Caddy 2.11.6 | Winner |
|---|---|---|---|
| 1 KB files over HTTPS | 79,169 req/s | 29,449 req/s | nginx, 2.7x |
| 100 KB files over HTTPS | 14,998 req/s | 12,785 req/s | nginx, 1.2x |
| New TLS handshakes | 3,531/s | 4,274/s | Caddy, 1.2x |
| Reverse proxy, stock config, 90 s | 627 req/s, 44,005 failed | 6,885 req/s, 74 failed | Caddy |
| Reverse proxy, with the fix, 90 s | 22,153 req/s | 18,048 req/s | nginx, 1.2x |
| p99 latency at 5,000 req/s | 0.68 ms | 2.52 ms | nginx |
| CPU at 5,000 req/s | 0.23 cores | 0.50 cores | nginx |
| RAM, idle and peak | 3.6 / 18.0 MiB | 10.0 / 44.8 MiB | nginx |
| HTTPS setup | Manual certificates | Automatic | Caddy |
nginx wins on raw speed, tail latency, CPU, and memory. Caddy wins on setup and on new handshakes, and its stock proxy config slows down under sustained load instead of failing outright. Both need one proxy fix, and with the fixes in place nginx leads by 1.2x, on a backend that capped nginx’s number. Pick Caddy for convenience: automatic HTTPS and a short config. Pick nginx for maximum throughput per core or the tightest tail latency. Whichever you run as a reverse proxy, apply its connection fix from this post. Your 2 AM self will thank you.
How we tested
One machine, everything in Docker, every part pinned to its own cores so the contestants do not fight over CPU.
- Hardware: Intel Core i7-11800H laptop (8 cores, 16 threads), 64 GB RAM, Ubuntu 22.04, kernel 6.8, Docker 29.8.
- CPU pinning with
--cpuset-cpus:- Server under test: 2 physical cores. Their hyperthread siblings stayed idle.
- Load generator (oha 1.16.0): 3 physical cores (6 threads).
- Backend for the proxy tests (traefik/whoami 1.11.0): 2 physical cores (4 threads).
- Versions: nginx 1.31.6 (current mainline,
nginx:1.31.6-alpine), nginx 1.28.3 for the older-version proxy tests, and Caddy 2.11.6 (caddy:2.11.6-alpine, the newest official Caddy image on October 5, 2026; the 2.11.7 binary shipped two days earlier). - Method: HTTP/1.1 everywhere, 100 concurrent connections, 20 seconds per test, 3 runs each, median reported. Every test starts with fresh containers, so leftover sockets from one test cannot hurt the next. Run-to-run spread stayed at or under 3.1% on the file tests and up to 7.5% on the proxy tests. A separate 90-second sustained proxy test runs once per config.
- TLS: the same self-signed ECDSA P-256 certificate on both servers, TLS 1.3.
- Counting: only
2xxresponses count as successful. oha’s own requests-per-second figure counts 502s too, which hides the failure this post is about. - Access logs off on both. Caddy logs nothing by default. nginx logs every request by default, so we turned that off to compare the servers instead of the log pipeline. We did not measure what the default nginx access log costs.
The kit runs the whole suite:
git clone https://github.com/KingPin/sumguy-examples.gitcd sumguy-examples/networking/nginx-vs-caddy-bench./bench.sh # about 45 minutes: 20 s per test, 3 runs, plus 90 s sustained testspython3 summarize.py # median tablesYour CPU layout differs from ours, so override the pinning. To repeat the older-nginx tests, swap the image and the config list:
SERVER_CPUS=0,1 LOADGEN_CPUS=2-5 BACKEND_CPUS=6,7 ./bench.shNGINX_IMAGE=nginx:1.28.3-alpine SERVERS="nginx nginx-upstream nginx-keepalive" OUT=results-1.28 ./bench.shThe limits, before you quote anything
- One machine. The load generator shares memory bandwidth and the kernel network stack with the server, even on separate cores.
- Docker bridge networking, not a real NIC.
- Other containers kept running on the box. A snapshot before testing showed them using about 0.17 cores in total.
- No HTTP/2, no HTTP/3, no compression. Only 1 KB and 100 KB files.
The server was CPU-bound (1.86 to 1.95 of 2 cores busy) in every max-throughput test except nginx proxying with working keepalive. There the backend was the bottleneck (whoami at 3.61 of 4 threads, nginx at 0.77 to 1.02 cores). So nginx’s real proxy ceiling is higher than the 22,120 we report. Our backend could not go faster.
The configs
Both are near stock: the official image, plus the minimum to serve a file, terminate TLS, and proxy one path. nginx uses worker_processes auto, and it started 2 workers inside the 2-core cpuset.
worker_processes auto;events { worker_connections 1024; }
http { include /etc/nginx/mime.types; access_log off; sendfile on; keepalive_timeout 65;
server { listen 80; listen 443 ssl; server_name bench.lab; ssl_certificate /certs/cert.pem; ssl_certificate_key /certs/key.pem;
root /srv/www;
location /proxy/ { proxy_pass http://whoami:80/; } }}That proxy_pass http://whoami:80/; line is the short form, with no upstream block. Hold that thought.
Caddy gets the same job. auto_https off is there only because the certificate is local, not from Let’s Encrypt. Everything else is default.
{ auto_https off}
(site) { handle_path /proxy/* { reverse_proxy whoami:80 } handle { root * /srv/www file_server }}
http://bench.lab { import site}
https://bench.lab { tls /certs/cert.pem /certs/key.pem import site}Results: serving files on 2 cores
Medians in successful requests per second, with p99 latency in milliseconds.
| Test | nginx 1.31.6 | Caddy 2.11.6 | Ratio |
|---|---|---|---|
| 1 KB file, HTTP | 92,609 (p99 1.4) | 20,981 (p99 14.0) | nginx 4.4x |
| 1 KB file, HTTPS | 79,169 (p99 1.9) | 29,449 (p99 11.2) | nginx 2.7x |
| 100 KB file, HTTP | 68,955 (p99 2.0) | 13,544 (p99 25.5) | nginx 5.1x |
| 100 KB file, HTTPS | 14,998 (p99 9.1) | 12,785 (p99 23.6) | nginx 1.2x |
| New TLS connection per request | 3,531 (p99 32.1) | 4,274 (p99 44.9) | Caddy 1.2x |
nginx 1.28.3 landed within 2% of 1.31.6 on every file test, so the nginx version does not matter here. Per core, nginx is the faster engine in every keep-alive test, and its p99 latency is 2.6x to 13x tighter.
Why the 100 KB gap shrinks over HTTPS
nginx serves plain HTTP files with sendfile(), which moves the file from the page cache to the socket inside the kernel. We checked with a pair of single 20-second runs: with sendfile off, nginx’s 100 KB HTTP result fell from 69,989 to 27,989 requests per second. With TLS, nginx has to read and encrypt the file in userspace (we did not configure kernel TLS), so its 100 KB HTTPS number is 14,998. Caddy was slow on 100 KB files over plain HTTP too (13,544), so adding TLS cost Caddy little, and the gap shrinks to 1.2x.
The Caddy plain-HTTP oddity
Caddy served 1 KB files faster over HTTPS (29,449) than over plain HTTP (20,981). That looks backwards, so we checked. A Caddy respond "hello" 200 route with no file ran 63,045 requests per second over HTTP and 58,161 over HTTPS. Plain HTTP itself is fine. The slow part is file serving over plain TCP.
A 5-second strace showed the two paths behave differently (24,914 write() and 20,945 nanosleep() calls on HTTP, 11,141 and 1,178 on HTTPS), but strace slows the server and we did not trace a root cause. Most homelabs serve Caddy over HTTPS anyway.
Handshakes: the one Caddy win
The handshake test runs oha with --disable-keepalive, so every request is a full new TLS 1.3 handshake. Caddy did 21% more handshakes per second. nginx had the tighter tail (p99 32.1 vs 44.9 ms). We did not measure why Caddy was faster, so we will not guess.
The proxy trap: 28,231 requests, then 502s
Here are the 20-second proxy numbers. “Failed” counts every non-2xx response across the 3 runs.
| Config | Successful req/s | p99 ms | Failed | Server cores |
|---|---|---|---|---|
nginx 1.31.6, proxy_pass http://whoami:80/ | 1,411 | 219.5 | 20,226 | 1.92 |
nginx 1.31.6, proxy_pass to an upstream block | 22,120+ | 23.3 | 0 | 1.02 |
Caddy 2.11.6, stock reverse_proxy | 12,496 | 51.7 | 0 | 1.88 |
| Caddy 2.11.6, 128 idle backend connections | 17,889 | 16.0 | 0 | 1.86 |
The 1,411 is misleading on its own. In every run, nginx served exactly 28,231 successful requests, and then it returned 502 Bad Gateway for everything else until the test ended. That is not a slow proxy. That is an outage with a counter on it.
What is happening
A proxy_pass pointed straight at a hostname gets no connection reuse. Every proxied request opens a new TCP connection to the backend, and nginx closes it afterward. The side that closes a TCP connection first keeps the socket in TIME_WAIT for 60 seconds. On Linux that 60 seconds is compiled into the kernel. The net.ipv4.tcp_tw_reuse setting can reuse those ports for new outgoing connections, but its default (2) applies to loopback only, and we left it at the default.
Every outgoing connection needs a local port. Inside the container, the ephemeral port range was 32768 to 60999, which is 28,232 ports. We counted TIME_WAIT sockets in the nginx container 6 seconds into a 10-second run: 28,158. With every port parked in TIME_WAIT, nginx cannot open a new backend connection, so it returns 502. One port short of the full range and 28,231 successful requests per run is no coincidence.
Why nginx 1.31 hits the wall harder than 1.28
nginx 1.29.7 (March 2026) changed the proxy defaults. The changelog says: “now ngx_http_proxy_module supports keepalive by default; the default value for “proxy_http_version” is “1.1”; the “Connection” proxy header is not sent by default anymore.” The keepalive directive in an upstream block is also on by default now, with 32 idle connections per worker.
That only helps backends defined in an upstream block. A bare proxy_pass http://whoami:80/ got no connection reuse in our tests on 1.31.6, and it also lost the old Connection: close header. We checked what the backend received:
| Header sent to backend | TIME_WAIT on nginx side | TIME_WAIT on backend side | All responses/s, incl. 502s | |
|---|---|---|---|---|
| nginx 1.28.3 | Connection: close | 14,031 | 3,438 | 12,487 |
| nginx 1.31.6 | none | 28,158 | 0 | 2,904 |
With Connection: close, the backend often closes first and keeps the TIME_WAIT socket on its side, which costs nginx no ports. Without it, nginx closes every connection and keeps every TIME_WAIT socket.
On 1.28.3, which side closes first depends on load. At full load for 90 seconds, nginx held a flat 14,000 or so TIME_WAIT sockets and never ran out (8,169 successful requests per second, zero failures). At a steady 1,000 requests per second, nginx held 9,354 TIME_WAIT sockets after 10 seconds (the backend held 4,155), filled the range by 40 seconds, and started returning 502s.
Here is a 2-minute run through the same bare proxy_pass config, with oha targeting a steady 1,000 requests per second:
| Seconds | 1.31.6 OK | 1.31.6 502 | 1.28.3 OK | 1.28.3 502 |
|---|---|---|---|---|
| 0 to 30 | 28,231 | 1,229 | 29,999 | 0 |
| 30 to 60 | 0 | 17,146 | 6,333 | 14,727 |
| 60 to 90 | 26,446 | 13,454 | 33,455 | 5,485 |
| 90 to 120 | 1,785 | 18,093 | 13,752 | 11,236 |
Both versions work for about half a minute, fail for about half a minute, then recover as the first TIME_WAIT sockets expire, and repeat. The app looks healthy, the backend is idle, and nginx serves 502s. That is the 2 AM page.
The math for your homelab
When the proxy closes every backend connection itself, as a bare proxy_pass did on nginx 1.31.6 (and by the changelog, on 1.29.7 and newer), the limit is 28,232 ports divided by 60 seconds: about 470 new backend connections per second, sustained, per backend address. Older nginx tolerates more because the backend absorbs part of the TIME_WAIT load, but our 1,000 requests per second run broke it anyway.
Most homelabs never sustain 470 requests per second to one backend. A scraper, a sync client stuck in a retry loop, or your own load test can.
The nginx fix
On nginx 1.29.7 or newer, define the backend in an upstream block. Keepalive is on by default there:
upstream app { server whoami:80;}
server { location /proxy/ { proxy_pass http://app/; }}On anything older, add the keepalive lines yourself. That covers the nginx packages in current Debian and Ubuntu releases: Debian 12 ships 1.22.1, Debian 13 ships 1.26.3, Ubuntu 24.04 ships 1.24.0, and Ubuntu 26.04 ships 1.28.3.
upstream app { server whoami:80; keepalive 32;}
server { location /proxy/ { proxy_pass http://app/; proxy_http_version 1.1; proxy_set_header Connection ""; }}keepalive 32 caches up to 32 idle backend connections per worker process. The other two lines are required: proxy_http_version 1.1 because old nginx speaks HTTP/1.0 to backends by default, and the empty Connection header because old nginx sends Connection: close by default. Miss either one and you keep the old behavior with no warning.
On 1.28.3, the upstream block without those lines did no better than a bare proxy_pass (10,096 vs 10,144 requests per second). With the three lines it hit 22,251, limited by our backend, with nginx at 0.77 cores, and 22,393 over the 90-second sustained test. The same lines also work on new nginx: a 20-second check on 1.31.6 ran 21,589 requests per second with zero failures. That makes the second config the safe one to copy.
Caddy has a smaller version of the same trap
Caddy’s reverse_proxy reuses backend connections, but it keeps at most 32 idle connections per backend (keepalive_idle_conns_per_host, default 32 in the Caddy docs). Our test runs 100 requests at once. The connections above 32 close after each request, and the closed ones sit in TIME_WAIT on Caddy’s side.
Twenty seconds is too short to show it. The 90-second sustained test does:
| 90 s at full load | Successful req/s | p99 ms | Failed | TIME_WAIT at 20 / 40 / 60 / 80 s |
|---|---|---|---|---|
nginx 1.31.6, bare proxy_pass | 627 | 259.6 | 44,005 | 28,231 / 28,231 / 28,231 / 28,231 |
nginx 1.31.6, upstream block | 22,153 | 23.4 | 0 | 402 / 1,072 / 2,192 / 2,419 |
| Caddy 2.11.6, stock | 6,885 | 200.2 | 74 | 23,358 / 28,183 / 28,184 / 25,126 |
| Caddy 2.11.6, 128 idle connections | 18,048 | 16.5 | 0 | 0 / 0 / 0 / 0 |
Stock Caddy filled the port range within 40 seconds. It did not fall over the way bare nginx did. It slowed to 6,885 successful requests per second while its TIME_WAIT sockets filled the range, and it returned 74 errors. We did not prove the full port range caused the whole slowdown. At a steady 1,000 requests per second for 2 minutes, stock Caddy held 22 TIME_WAIT sockets and served every request. Over 32 concurrent requests is the condition for the leak, and the leak has to pass about 470 closed connections per second for a minute before the ports run out. We only tested 100 concurrent requests, so we cannot say where between 33 and 100 that starts.
The fix is one setting. Make the idle pool at least as large as your peak concurrency to that backend:
reverse_proxy whoami:80 { transport http { keepalive_idle_conns_per_host 128 }}With it, Caddy proxied 17,889 requests per second in the 20-second test (up from 12,496), its p99 dropped from 51.7 ms to 16.0 ms, and the sustained test held 18,048 with zero TIME_WAIT sockets. That is within 20% of nginx’s number, and nginx’s number was capped by our backend.
Tail latency at a fixed rate
Max-throughput tests show the ceiling. We also ran a fixed 5,000 requests per second through the HTTPS proxy for 20 seconds, using oha’s --latency-correction so queueing delay counts against the server the way a real client feels it.
| p50 ms | p99 ms | p99.9 ms | Server cores | Failed | |
|---|---|---|---|---|---|
nginx 1.31.6, upstream block | 0.28 | 0.68 | 30.30 | 0.23 | 0 |
| nginx 1.28.3, keepalive lines | 0.26 | 0.64 | 29.08 | 0.18 | 0 |
| Caddy 2.11.6, stock | 0.36 | 2.52 | 49.43 | 0.50 | 0 |
| Caddy 2.11.6, 128 idle connections | 0.37 | 2.87 | 76.39 | 0.50 | 0 |
nginx 1.31.6, bare proxy_pass | 941 | 12,680 | 13,071 | 1.73 | 18,651 |
With working keepalive, nginx used under half the CPU Caddy did for the same 5,000 requests per second. At this rate Caddy’s bigger idle pool made no useful difference (its p99.9 was 76 ms vs 49 ms, which we put down to noise).
Caddy’s p95 stayed under 1 ms. Its p99 moved around: 1.9 to 2.9 ms in the final runs above, up to 46.7 ms in one earlier session on the same box and 33.2 ms in another. nginx with keepalive stayed between 0.59 and 0.74 ms at p99 in every session. Both servers show 25 ms or more at p99.9, so part of that far tail is the box itself. We have not traced the cause of Caddy’s p99 swings.
What this means for your homelab
- Grep your nginx configs for bare
proxy_passlines today. Anything that proxies tohttp://host:portinstead of an upstream name gets no connection reuse. Move it into anupstreamblock with the keepalive lines. It removes the 502 cliff at about 470 new connections per second. - Caddy users: you are fine at homelab load. Stock Caddy served a steady 1,000 requests per second for 2 minutes with only 22 TIME_WAIT sockets. If you push more than 32 concurrent requests at one backend for minutes at a time, raise
keepalive_idle_conns_per_host. - Starting fresh, pick Caddy. Automatic HTTPS and a shorter config. Its proxy trap needs far more load to trigger than nginx’s.
- Already happy on nginx? Stay. Fix the proxy config and you have the faster server per core with the tighter tail.
- Pick nginx when the numbers matter. High throughput per core, tight p99, or a small shared box where 0.23 cores vs 0.50 cores at 5,000 requests per second adds up.
- Do not choose on RAM. Caddy peaked at 44.8 MiB to nginx’s 18.0, measured from the container’s cgroup with page cache included. Choosing between these two on RAM is like weighing a forklift against a golf cart to save a parking space.
What we did not test
- HTTP/2 and HTTP/3
- Compression (gzip, brotli, zstd)
- nginx with its default access log on
- A real NIC instead of Docker bridge networking
- More than 2 server cores, where the per-core picture could shift
- File sizes other than 1 KB and 100 KB
- The cause of Caddy’s plain-HTTP file slowness and its p99 swings
The kit is the point. Run it on your hardware with your CPU layout and compare. If your numbers disagree with ours, open an issue on the examples repo.
Common Questions
Does nginx proxy_pass use keepalive by default?
Not for a bare proxy_pass http://host:port, which got no connection reuse in our tests on nginx 1.31.6. Since nginx 1.29.7, a proxy_pass to an upstream block reuses backend connections by default. On nginx 1.28 and older, upstream keepalive needs keepalive in the upstream block, proxy_http_version 1.1, and an empty Connection header.
Why does nginx return 502 Bad Gateway under load?
One common cause is local port exhaustion between nginx and the backend. Without upstream keepalive, nginx opens a new connection per request, and each closed connection holds a port in TIME_WAIT for 60 seconds. Above about 470 new connections per second, the 28,232 ports in the default range run out and nginx returns 502 until they free up.
Does Caddy reverse_proxy reuse backend connections?
Yes, up to 32 idle connections per backend by default. Caddy closes connections beyond that pool after each request, so sustained load above 32 concurrent requests to one backend leaks ports into TIME_WAIT. At 100 concurrent requests for 90 seconds, stock Caddy slowed to 6,885 requests per second. Setting keepalive_idle_conns_per_host 128 restored 18,048.
Should I switch from nginx to Caddy for performance?
No. nginx was faster per core on every keep-alive file test we ran, by 1.2x to 5.1x, so switching to Caddy will not speed anything up. Switch for convenience: automatic HTTPS and a shorter config. If your nginx proxy throws 502s under load, add upstream keepalive before you migrate anything.