Skip to content
Go back

nginx vs Caddy: Measured on 2 Cores

By KingPin 19 min read
nginx vs Caddy: Measured on 2 Cores
Contents

The short version: nginx is faster, and that is not the story

We pinned nginx and Caddy to the same two CPU cores and hit them with the same load. nginx is faster per core on every keep-alive file test, by 1.2x to 5.1x. At homelab traffic, that gap does not matter. Caddy on 2 cores still served 21,000 to 29,000 small files per second, and your Jellyfin box does not see that on its best day.

The results that matter are about connections to the backend, and both servers ship a default worth changing. The common nginx form, proxy_pass pointed straight at a backend hostname, opens a new TCP connection for every request. Under sustained load it runs out of local ports and answers with 502 Bad Gateway. On nginx 1.31.6, the first 28,231 requests worked and everything after that failed. Caddy reuses backend connections, but it keeps only 32 idle ones per backend. Past 32 concurrent requests it slowly leaks ports too. Over 90 seconds at 100 concurrent requests, its TIME_WAIT sockets filled the port range and its throughput dropped to 6,885 requests per second. Each fix is a few lines, shown below.

Full example: Clone the working files at github.com/KingPin/sumguy-examples/networking/nginx-vs-caddy-bench

This is the first post in SumGuy’s Lab: posts built on numbers we measured on our own hardware, with the test kit published so you can check our work. The box at the top of this page lists the rig and its limits.

The scorecard

Test (2 cores)nginx 1.31.6Caddy 2.11.6Winner
1 KB files over HTTPS79,169 req/s29,449 req/snginx, 2.7x
100 KB files over HTTPS14,998 req/s12,785 req/snginx, 1.2x
New TLS handshakes3,531/s4,274/sCaddy, 1.2x
Reverse proxy, stock config, 90 s627 req/s, 44,005 failed6,885 req/s, 74 failedCaddy
Reverse proxy, with the fix, 90 s22,153 req/s18,048 req/snginx, 1.2x
p99 latency at 5,000 req/s0.68 ms2.52 msnginx
CPU at 5,000 req/s0.23 cores0.50 coresnginx
RAM, idle and peak3.6 / 18.0 MiB10.0 / 44.8 MiBnginx
HTTPS setupManual certificatesAutomaticCaddy

nginx wins on raw speed, tail latency, CPU, and memory. Caddy wins on setup and on new handshakes, and its stock proxy config slows down under sustained load instead of failing outright. Both need one proxy fix, and with the fixes in place nginx leads by 1.2x, on a backend that capped nginx’s number. Pick Caddy for convenience: automatic HTTPS and a short config. Pick nginx for maximum throughput per core or the tightest tail latency. Whichever you run as a reverse proxy, apply its connection fix from this post. Your 2 AM self will thank you.

How we tested

One machine, everything in Docker, every part pinned to its own cores so the contestants do not fight over CPU.

The kit runs the whole suite:

Terminal window
git clone https://github.com/KingPin/sumguy-examples.git
cd sumguy-examples/networking/nginx-vs-caddy-bench
./bench.sh # about 45 minutes: 20 s per test, 3 runs, plus 90 s sustained tests
python3 summarize.py # median tables

Your CPU layout differs from ours, so override the pinning. To repeat the older-nginx tests, swap the image and the config list:

Terminal window
SERVER_CPUS=0,1 LOADGEN_CPUS=2-5 BACKEND_CPUS=6,7 ./bench.sh
NGINX_IMAGE=nginx:1.28.3-alpine SERVERS="nginx nginx-upstream nginx-keepalive" OUT=results-1.28 ./bench.sh

The limits, before you quote anything

The server was CPU-bound (1.86 to 1.95 of 2 cores busy) in every max-throughput test except nginx proxying with working keepalive. There the backend was the bottleneck (whoami at 3.61 of 4 threads, nginx at 0.77 to 1.02 cores). So nginx’s real proxy ceiling is higher than the 22,120 we report. Our backend could not go faster.

The configs

Both are near stock: the official image, plus the minimum to serve a file, terminate TLS, and proxy one path. nginx uses worker_processes auto, and it started 2 workers inside the 2-core cpuset.

nginx.conf
worker_processes auto;
events { worker_connections 1024; }
http {
include /etc/nginx/mime.types;
access_log off;
sendfile on;
keepalive_timeout 65;
server {
listen 80;
listen 443 ssl;
server_name bench.lab;
ssl_certificate /certs/cert.pem;
ssl_certificate_key /certs/key.pem;
root /srv/www;
location /proxy/ {
proxy_pass http://whoami:80/;
}
}
}

That proxy_pass http://whoami:80/; line is the short form, with no upstream block. Hold that thought.

Caddy gets the same job. auto_https off is there only because the certificate is local, not from Let’s Encrypt. Everything else is default.

Caddyfile
{
auto_https off
}
(site) {
handle_path /proxy/* {
reverse_proxy whoami:80
}
handle {
root * /srv/www
file_server
}
}
http://bench.lab {
import site
}
https://bench.lab {
tls /certs/cert.pem /certs/key.pem
import site
}

Results: serving files on 2 cores

Medians in successful requests per second, with p99 latency in milliseconds.

Testnginx 1.31.6Caddy 2.11.6Ratio
1 KB file, HTTP92,609 (p99 1.4)20,981 (p99 14.0)nginx 4.4x
1 KB file, HTTPS79,169 (p99 1.9)29,449 (p99 11.2)nginx 2.7x
100 KB file, HTTP68,955 (p99 2.0)13,544 (p99 25.5)nginx 5.1x
100 KB file, HTTPS14,998 (p99 9.1)12,785 (p99 23.6)nginx 1.2x
New TLS connection per request3,531 (p99 32.1)4,274 (p99 44.9)Caddy 1.2x

nginx 1.28.3 landed within 2% of 1.31.6 on every file test, so the nginx version does not matter here. Per core, nginx is the faster engine in every keep-alive test, and its p99 latency is 2.6x to 13x tighter.

Why the 100 KB gap shrinks over HTTPS

nginx serves plain HTTP files with sendfile(), which moves the file from the page cache to the socket inside the kernel. We checked with a pair of single 20-second runs: with sendfile off, nginx’s 100 KB HTTP result fell from 69,989 to 27,989 requests per second. With TLS, nginx has to read and encrypt the file in userspace (we did not configure kernel TLS), so its 100 KB HTTPS number is 14,998. Caddy was slow on 100 KB files over plain HTTP too (13,544), so adding TLS cost Caddy little, and the gap shrinks to 1.2x.

The Caddy plain-HTTP oddity

Caddy served 1 KB files faster over HTTPS (29,449) than over plain HTTP (20,981). That looks backwards, so we checked. A Caddy respond "hello" 200 route with no file ran 63,045 requests per second over HTTP and 58,161 over HTTPS. Plain HTTP itself is fine. The slow part is file serving over plain TCP.

A 5-second strace showed the two paths behave differently (24,914 write() and 20,945 nanosleep() calls on HTTP, 11,141 and 1,178 on HTTPS), but strace slows the server and we did not trace a root cause. Most homelabs serve Caddy over HTTPS anyway.

Handshakes: the one Caddy win

The handshake test runs oha with --disable-keepalive, so every request is a full new TLS 1.3 handshake. Caddy did 21% more handshakes per second. nginx had the tighter tail (p99 32.1 vs 44.9 ms). We did not measure why Caddy was faster, so we will not guess.

The proxy trap: 28,231 requests, then 502s

Here are the 20-second proxy numbers. “Failed” counts every non-2xx response across the 3 runs.

ConfigSuccessful req/sp99 msFailedServer cores
nginx 1.31.6, proxy_pass http://whoami:80/1,411219.520,2261.92
nginx 1.31.6, proxy_pass to an upstream block22,120+23.301.02
Caddy 2.11.6, stock reverse_proxy12,49651.701.88
Caddy 2.11.6, 128 idle backend connections17,88916.001.86

The 1,411 is misleading on its own. In every run, nginx served exactly 28,231 successful requests, and then it returned 502 Bad Gateway for everything else until the test ended. That is not a slow proxy. That is an outage with a counter on it.

What is happening

A proxy_pass pointed straight at a hostname gets no connection reuse. Every proxied request opens a new TCP connection to the backend, and nginx closes it afterward. The side that closes a TCP connection first keeps the socket in TIME_WAIT for 60 seconds. On Linux that 60 seconds is compiled into the kernel. The net.ipv4.tcp_tw_reuse setting can reuse those ports for new outgoing connections, but its default (2) applies to loopback only, and we left it at the default.

Every outgoing connection needs a local port. Inside the container, the ephemeral port range was 32768 to 60999, which is 28,232 ports. We counted TIME_WAIT sockets in the nginx container 6 seconds into a 10-second run: 28,158. With every port parked in TIME_WAIT, nginx cannot open a new backend connection, so it returns 502. One port short of the full range and 28,231 successful requests per run is no coincidence.

Why nginx 1.31 hits the wall harder than 1.28

nginx 1.29.7 (March 2026) changed the proxy defaults. The changelog says: “now ngx_http_proxy_module supports keepalive by default; the default value for “proxy_http_version” is “1.1”; the “Connection” proxy header is not sent by default anymore.” The keepalive directive in an upstream block is also on by default now, with 32 idle connections per worker.

That only helps backends defined in an upstream block. A bare proxy_pass http://whoami:80/ got no connection reuse in our tests on 1.31.6, and it also lost the old Connection: close header. We checked what the backend received:

Header sent to backendTIME_WAIT on nginx sideTIME_WAIT on backend sideAll responses/s, incl. 502s
nginx 1.28.3Connection: close14,0313,43812,487
nginx 1.31.6none28,15802,904

With Connection: close, the backend often closes first and keeps the TIME_WAIT socket on its side, which costs nginx no ports. Without it, nginx closes every connection and keeps every TIME_WAIT socket.

On 1.28.3, which side closes first depends on load. At full load for 90 seconds, nginx held a flat 14,000 or so TIME_WAIT sockets and never ran out (8,169 successful requests per second, zero failures). At a steady 1,000 requests per second, nginx held 9,354 TIME_WAIT sockets after 10 seconds (the backend held 4,155), filled the range by 40 seconds, and started returning 502s.

Here is a 2-minute run through the same bare proxy_pass config, with oha targeting a steady 1,000 requests per second:

Seconds1.31.6 OK1.31.6 5021.28.3 OK1.28.3 502
0 to 3028,2311,22929,9990
30 to 60017,1466,33314,727
60 to 9026,44613,45433,4555,485
90 to 1201,78518,09313,75211,236

Both versions work for about half a minute, fail for about half a minute, then recover as the first TIME_WAIT sockets expire, and repeat. The app looks healthy, the backend is idle, and nginx serves 502s. That is the 2 AM page.

The math for your homelab

When the proxy closes every backend connection itself, as a bare proxy_pass did on nginx 1.31.6 (and by the changelog, on 1.29.7 and newer), the limit is 28,232 ports divided by 60 seconds: about 470 new backend connections per second, sustained, per backend address. Older nginx tolerates more because the backend absorbs part of the TIME_WAIT load, but our 1,000 requests per second run broke it anyway.

Most homelabs never sustain 470 requests per second to one backend. A scraper, a sync client stuck in a retry loop, or your own load test can.

The nginx fix

On nginx 1.29.7 or newer, define the backend in an upstream block. Keepalive is on by default there:

nginx 1.29.7 and newer
upstream app {
server whoami:80;
}
server {
location /proxy/ {
proxy_pass http://app/;
}
}

On anything older, add the keepalive lines yourself. That covers the nginx packages in current Debian and Ubuntu releases: Debian 12 ships 1.22.1, Debian 13 ships 1.26.3, Ubuntu 24.04 ships 1.24.0, and Ubuntu 26.04 ships 1.28.3.

nginx 1.28 and older
upstream app {
server whoami:80;
keepalive 32;
}
server {
location /proxy/ {
proxy_pass http://app/;
proxy_http_version 1.1;
proxy_set_header Connection "";
}
}

keepalive 32 caches up to 32 idle backend connections per worker process. The other two lines are required: proxy_http_version 1.1 because old nginx speaks HTTP/1.0 to backends by default, and the empty Connection header because old nginx sends Connection: close by default. Miss either one and you keep the old behavior with no warning.

On 1.28.3, the upstream block without those lines did no better than a bare proxy_pass (10,096 vs 10,144 requests per second). With the three lines it hit 22,251, limited by our backend, with nginx at 0.77 cores, and 22,393 over the 90-second sustained test. The same lines also work on new nginx: a 20-second check on 1.31.6 ran 21,589 requests per second with zero failures. That makes the second config the safe one to copy.

Caddy has a smaller version of the same trap

Caddy’s reverse_proxy reuses backend connections, but it keeps at most 32 idle connections per backend (keepalive_idle_conns_per_host, default 32 in the Caddy docs). Our test runs 100 requests at once. The connections above 32 close after each request, and the closed ones sit in TIME_WAIT on Caddy’s side.

Twenty seconds is too short to show it. The 90-second sustained test does:

90 s at full loadSuccessful req/sp99 msFailedTIME_WAIT at 20 / 40 / 60 / 80 s
nginx 1.31.6, bare proxy_pass627259.644,00528,231 / 28,231 / 28,231 / 28,231
nginx 1.31.6, upstream block22,15323.40402 / 1,072 / 2,192 / 2,419
Caddy 2.11.6, stock6,885200.27423,358 / 28,183 / 28,184 / 25,126
Caddy 2.11.6, 128 idle connections18,04816.500 / 0 / 0 / 0

Stock Caddy filled the port range within 40 seconds. It did not fall over the way bare nginx did. It slowed to 6,885 successful requests per second while its TIME_WAIT sockets filled the range, and it returned 74 errors. We did not prove the full port range caused the whole slowdown. At a steady 1,000 requests per second for 2 minutes, stock Caddy held 22 TIME_WAIT sockets and served every request. Over 32 concurrent requests is the condition for the leak, and the leak has to pass about 470 closed connections per second for a minute before the ports run out. We only tested 100 concurrent requests, so we cannot say where between 33 and 100 that starts.

The fix is one setting. Make the idle pool at least as large as your peak concurrency to that backend:

Caddyfile
reverse_proxy whoami:80 {
transport http {
keepalive_idle_conns_per_host 128
}
}

With it, Caddy proxied 17,889 requests per second in the 20-second test (up from 12,496), its p99 dropped from 51.7 ms to 16.0 ms, and the sustained test held 18,048 with zero TIME_WAIT sockets. That is within 20% of nginx’s number, and nginx’s number was capped by our backend.

Tail latency at a fixed rate

Max-throughput tests show the ceiling. We also ran a fixed 5,000 requests per second through the HTTPS proxy for 20 seconds, using oha’s --latency-correction so queueing delay counts against the server the way a real client feels it.

p50 msp99 msp99.9 msServer coresFailed
nginx 1.31.6, upstream block0.280.6830.300.230
nginx 1.28.3, keepalive lines0.260.6429.080.180
Caddy 2.11.6, stock0.362.5249.430.500
Caddy 2.11.6, 128 idle connections0.372.8776.390.500
nginx 1.31.6, bare proxy_pass94112,68013,0711.7318,651

With working keepalive, nginx used under half the CPU Caddy did for the same 5,000 requests per second. At this rate Caddy’s bigger idle pool made no useful difference (its p99.9 was 76 ms vs 49 ms, which we put down to noise).

Caddy’s p95 stayed under 1 ms. Its p99 moved around: 1.9 to 2.9 ms in the final runs above, up to 46.7 ms in one earlier session on the same box and 33.2 ms in another. nginx with keepalive stayed between 0.59 and 0.74 ms at p99 in every session. Both servers show 25 ms or more at p99.9, so part of that far tail is the box itself. We have not traced the cause of Caddy’s p99 swings.

What this means for your homelab

  1. Grep your nginx configs for bare proxy_pass lines today. Anything that proxies to http://host:port instead of an upstream name gets no connection reuse. Move it into an upstream block with the keepalive lines. It removes the 502 cliff at about 470 new connections per second.
  2. Caddy users: you are fine at homelab load. Stock Caddy served a steady 1,000 requests per second for 2 minutes with only 22 TIME_WAIT sockets. If you push more than 32 concurrent requests at one backend for minutes at a time, raise keepalive_idle_conns_per_host.
  3. Starting fresh, pick Caddy. Automatic HTTPS and a shorter config. Its proxy trap needs far more load to trigger than nginx’s.
  4. Already happy on nginx? Stay. Fix the proxy config and you have the faster server per core with the tighter tail.
  5. Pick nginx when the numbers matter. High throughput per core, tight p99, or a small shared box where 0.23 cores vs 0.50 cores at 5,000 requests per second adds up.
  6. Do not choose on RAM. Caddy peaked at 44.8 MiB to nginx’s 18.0, measured from the container’s cgroup with page cache included. Choosing between these two on RAM is like weighing a forklift against a golf cart to save a parking space.

What we did not test

The kit is the point. Run it on your hardware with your CPU layout and compare. If your numbers disagree with ours, open an issue on the examples repo.

Common Questions

Does nginx proxy_pass use keepalive by default?

Not for a bare proxy_pass http://host:port, which got no connection reuse in our tests on nginx 1.31.6. Since nginx 1.29.7, a proxy_pass to an upstream block reuses backend connections by default. On nginx 1.28 and older, upstream keepalive needs keepalive in the upstream block, proxy_http_version 1.1, and an empty Connection header.

Why does nginx return 502 Bad Gateway under load?

One common cause is local port exhaustion between nginx and the backend. Without upstream keepalive, nginx opens a new connection per request, and each closed connection holds a port in TIME_WAIT for 60 seconds. Above about 470 new connections per second, the 28,232 ports in the default range run out and nginx returns 502 until they free up.

Does Caddy reverse_proxy reuse backend connections?

Yes, up to 32 idle connections per backend by default. Caddy closes connections beyond that pool after each request, so sustained load above 32 concurrent requests to one backend leaks ports into TIME_WAIT. At 100 concurrent requests for 90 seconds, stock Caddy slowed to 6,885 requests per second. Setting keepalive_idle_conns_per_host 128 restored 18,048.

Should I switch from nginx to Caddy for performance?

No. nginx was faster per core on every keep-alive file test we ran, by 1.2x to 5.1x, so switching to Caddy will not speed anything up. Switch for convenience: automatic HTTPS and a shorter config. If your nginx proxy throws 502s under load, add upstream keepalive before you migrate anything.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Previous Post
Crawl4AI: Feed Your RAG From the Web
Next Post
Solar Plus Battery for Your Home Lab

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts