Skip to content
Go back

Docker Build Cache in CI: Make It Hit

By KingPin 14 min read
Docker Build Cache in CI: Make It Hit
Contents

Your Laptop Is Fast and Your CI Is a Goldfish

Locally, docker build finishes in four seconds because BuildKit remembers every layer. In CI, the same Dockerfile takes nine minutes because the runner is a fresh VM. It has no memory of yesterday, and the cache sat on a machine that no longer exists.

The fix is to store the cache somewhere that outlives the runner. My picks: use the registry cache backend with mode=max as the default for most people. Use type=gha if you live on GitHub-hosted runners, and always set a scope. If you self-host runners, skip both and keep a long-lived builder around. Then spend the rest of your time on the part that actually bites: cache that is “there” but still misses.

Why Ephemeral Runners Start Cold

BuildKit keeps its cache in its own storage on the builder. A GitHub-hosted runner, a fresh GitLab job container, or a Forgejo Actions job gets thrown away when the job ends. The cache goes with it. The next job starts from zero and rebuilds every layer.

You have three ways out:

  1. Export the cache to a registry (type=registry).
  2. Export the cache to the GitHub Actions cache service (type=gha).
  3. Stop throwing the builder away (self-hosted, persistent).

Options 1 and 2 need a builder that can export cache. This is where the driver matters.

Drivers: Which Builder Can Export Cache

Docker’s docs say the default docker driver supports the inline, local, registry, and gha backends, but only when the containerd image store is enabled. On a stock install without it, you need a different driver. The safe move is to create a docker-container builder, which is what docker/setup-buildx-action does for you.

Terminal window
docker buildx create --name ci-builder --driver docker-container --use
docker buildx inspect --bootstrap

If you see an error that your driver does not support the cache export type, this is why. Create the builder and move on.

Cache Modes: min vs max

The mode parameter on --cache-to decides which layers get exported. Both the registry and gha backends default to min.

The cost of max is more bytes to store and upload. The cost of min on a multi-stage build is that your expensive builder stage, the one with the compiler and npm ci, is not in the cache at all. Every run rebuilds it, and then you wonder why the cache “does nothing”.

For a multi-stage Dockerfile, use mode=max. For a single-stage image where everything lands in the final image, min is fine and cheaper.

Option 1: Registry Cache

The registry backend pushes the cache as a separate image. It works on anything that can reach a registry: GitHub Actions, GitLab CI, Forgejo and Gitea Actions, Jenkins, a shell script on a cron job.

The one rule from the docs: ref can be any valid name as long as it differs from the image’s push target. A dedicated tag is the usual choice.

Terminal window
docker buildx build \
--push -t ghcr.io/youruser/yourapp:latest \
--cache-from type=registry,ref=ghcr.io/youruser/yourapp:buildcache \
--cache-to type=registry,ref=ghcr.io/youruser/yourapp:buildcache,mode=max \
.

That covers it, and it is plain CLI, so it works in any CI. If the cache image does not exist yet (first run), the import step fails but the build continues. Do not panic at the first error line.

Here is the GitHub Actions version:

.github/workflows/build.yml
name: build
on:
push:
branches: [main]
jobs:
image:
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
steps:
- uses: actions/checkout@v7
- uses: docker/setup-buildx-action@v4
- uses: docker/login-action@v4
with:
registry: ghcr.io
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- uses: docker/build-push-action@v7
with:
context: .
push: true
tags: ghcr.io/youruser/yourapp:latest
cache-from: type=registry,ref=ghcr.io/youruser/yourapp:buildcache
cache-to: type=registry,ref=ghcr.io/youruser/yourapp:buildcache,mode=max

Registries That Choke on the Cache Image

Some registries do not support image indexes. Amazon ECR is the example in Docker’s docs. For those, the docs say to set image-manifest=true, which makes the exporter emit an image manifest instead of an index. Since BuildKit v0.21 this is already the default, so on a current builder you probably never touch it. On an older builder pointed at ECR, add it:

Terminal window
--cache-to type=registry,ref=<registry>/yourapp:buildcache,mode=max,image-manifest=true

oci-mediatypes defaults to true as well. Flip it to false only if your registry rejects OCI media types.

Also note ignore-error=true on cache-to if you would rather a flaky registry not fail the whole build. It defaults to false.

Option 2: The GitHub Actions Cache Backend

type=gha stores cache blobs in the GitHub Actions cache service. No extra registry tag, no extra credentials. It only works inside a GitHub Actions workflow, because the URL and token variables only exist there.

.github/workflows/build-gha.yml
name: build-gha
on:
push:
branches: [main]
jobs:
image:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: docker/setup-buildx-action@v4
- uses: docker/build-push-action@v7
with:
context: .
push: false
tags: yourapp:ci
cache-from: type=gha,scope=yourapp
cache-to: type=gha,scope=yourapp,mode=max

docker/setup-buildx-action gives you a buildx builder that can export cache, so keep it in the workflow.

Scope: The Reason Your Cache Keeps Vanishing

The default scope is buildkit. Per the docs, each build overwrites the previous build’s cache, leaving only the final one. If you build two images, or run a matrix, and every job uses the default scope, they take turns wiping each other. Each build may import a cache written by another image.

Give every image its own scope:

matrix with per-image scope
strategy:
matrix:
app: [api, worker, web]
steps:
- uses: docker/build-push-action@v7
with:
context: ./${{ matrix.app }}
cache-from: type=gha,scope=${{ matrix.app }}
cache-to: type=gha,scope=${{ matrix.app }},mode=max

Add the platform to the scope if you build several (scope=api-arm64). A different platform is a different cache anyway.

Limits, Eviction, and the v2 API

GitHub’s docs state a default limit of 10 GB per repository, and entries not accessed in over 7 days are removed. When the repo hits its maximum, GitHub saves the new cache and evicts older ones in order of last access date. Repository admins and owners can raise the limit, with extra storage billed (check GitHub’s current docs for the terms). With mode=max, a few big images can eat 10 GB fast. Then the eviction fires and you get cache thrashing: caches created and deleted constantly.

Two more gotchas from Docker’s docs:

If exports fail for access reasons, ignore-error=true on cache-to keeps the build green. It suppresses all export errors though, so you will not notice a broken cache. Use it on purpose.

gha vs registry in One Paragraph

Registry cache is portable, has no size cap besides your registry’s, and survives a move to a different CI. The gha backend is zero-setup on GitHub-hosted runners but lives under a 10 GB repo budget shared with everything else that uses actions/cache. If you only ever run on GitHub, gha with a scope is fine. If you might leave, or you build large images, use the registry.

Option 3: The Persistent Builder (Home Lab Mode)

If you run your own runner, a long-lived buildx builder is the simplest fix. The docker-container builder keeps its cache in a Docker volume, and the volume survives between jobs. No export, no import, no network round trip.

Terminal window
# once, on the runner host
docker buildx create --name homelab --driver docker-container
Terminal window
# in every job on that runner
docker buildx build --builder homelab --load -t yourapp:ci .

Do not create a new builder per job, or you have rebuilt the goldfish. Prune now and then with docker buildx prune so the volume does not eat the disk. Using a registry cache for a runner that never dies is like hiring a forklift to move a couch. Technically it works, but your neighbors will have questions.

Inline Cache and Its Limits

type=inline embeds the cache inside the pushed image itself. It is the least setup: push the image, import from the same ref next time.

Terminal window
docker buildx build --push -t ghcr.io/youruser/yourapp:latest \
--cache-to type=inline \
--cache-from type=registry,ref=ghcr.io/youruser/yourapp:latest \
.

Docker’s docs list the catches. Inline cache supports only min mode, takes no extra parameters, and works only with the image exporter. It also does not separate your output from your cache. On multi-stage builds, that means your builder stages are never cached. Use inline for a single-stage image you already push. Otherwise use the registry backend.

Cache Mounts Do Not Ride Along

RUN --mount=type=cache,target=/root/.cache/pip is great locally. The docs describe cache mounts as persisting across builds on the same builder. They do not mention external backends, and Docker’s GitHub Actions page says BuildKit does not keep cache mounts in the GitHub cache by default. On an ephemeral runner, a cache mount starts empty on every run, and --cache-to does not change that.

The community workaround that Docker’s own docs point to is reproducible-containers/buildkit-cache-dance. It saves and injects the mount directory through actions/cache. As of this writing (October 2026) its latest release is v3.4.0, the repo is not archived, and it was last pushed in August 2026. Docker’s docs show actions/cache saving the directory, and the dance action injecting it into the build. Read its README for the current inputs, since they have changed between major versions.

On a persistent builder you skip all this, because the mount lives on the builder.

Why the Cache Is There but Still Misses

You import the cache, the log says it loaded, and every layer still rebuilds. This is almost always your Dockerfile, not your CI.

Bad Ordering

If you copy the whole source tree before installing dependencies, any code edit busts the dependency layer.

Dockerfile (cache-hostile)
FROM node:22-slim
WORKDIR /app
COPY . .
RUN npm ci
RUN npm run build

Copy the manifest files first, install, then copy the rest:

Dockerfile (cache-friendly)
FROM node:22-slim AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build
FROM node:22-slim
WORKDIR /app
COPY --from=build /app/dist ./dist
CMD ["node", "dist/server.js"]

Now npm ci reruns only when the lockfile changes. And this is a multi-stage build, so remember mode=max, or the build stage never reaches the cache.

Build ARGs That Change Every Run

Docker’s docs say plainly that build arguments cause cache invalidation. A --build-arg BUILD_DATE=$(date) or GIT_SHA that changes every commit will bust every instruction that sees it. Declare such ARGs as late as possible, right before the instruction that needs them, ideally in the final stage.

FROM node:22-slim AS build
ARG GIT_SHA
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build
ARG GIT_SHA
LABEL org.opencontainers.image.revision=$GIT_SHA

File Metadata, Not Timestamps

For COPY and ADD, Docker’s docs say the builder calculates the cache checksum from file metadata, and that the modification time (mtime) is not taken into account. So git checkout giving every file a fresh timestamp does not bust the cache. Changed content and changed metadata such as permissions do. I did not find a doc that lists exactly which metadata fields are hashed, so I will not guess. If a file mode flips between environments (a script that is executable on one checkout and not another), expect a miss.

A Missing .dockerignore

Without a .dockerignore, the build context includes .git, node_modules, local logs, and whatever your CI wrote into the workspace. Any change in there busts a COPY . . layer. A short file fixes it:

.dockerignore
.git
node_modules
dist
*.log
.github

The Moving Base Image

FROM node:22-slim is a tag, and tags move. When the maintainers push a new build, your first FROM resolves to a new digest, and every layer after it misses. This is the correct behavior, since you want security patches. It also explains the Monday morning full rebuild. If you want stable cache, pin a digest and update it on purpose with Renovate or Dependabot.

Other Quiet Misses

How to See Whether It Hit

Run with plain progress output so nothing is hidden:

Terminal window
docker buildx build --progress=plain \
--cache-from type=registry,ref=ghcr.io/youruser/yourapp:buildcache \
-t yourapp:ci .

A healthy run looks like this (illustrative, trimmed):

#1 [internal] load build definition from Dockerfile
#4 importing cache manifest from ghcr.io/youruser/yourapp:buildcache
#4 DONE 0.8s
#7 [build 3/6] COPY package.json package-lock.json ./
#7 CACHED
#8 [build 4/6] RUN npm ci
#8 CACHED
#9 [build 5/6] COPY . .
#9 DONE 0.3s
#10 [build 6/6] RUN npm run build
#10 DONE 12.4s

Read it top to bottom:

  1. The importing cache manifest from line shows the import ran. If it is missing, your --cache-from never reached BuildKit. If it fails, the cache ref does not exist or the login failed.
  2. CACHED lines are hits. The first layer without it is where the chain broke.
  3. Everything after the first miss rebuilds. Find the first miss, and look at the instruction and the files above it.

To inspect what a builder holds locally, use docker buildx du. On an ephemeral runner this only shows you the current job, so it is mostly useful on a persistent builder.

The SumGuy Take

Pick by where your runners live:

  1. Self-hosted runner: persistent docker-container builder. Done.
  2. GitHub-hosted, small images, staying on GitHub: type=gha, mode=max, and a scope per image.
  3. Anything else, big images, or you might switch CI: registry cache, mode=max, a dedicated :buildcache tag.

Then fix your Dockerfile order and write the .dockerignore. Your 2 AM self will appreciate it when the hotfix build takes ninety seconds instead of nine minutes.

Common Questions

Does Docker cache mount work in GitHub Actions?

No, not through --cache-to. Cache mounts from RUN --mount=type=cache are not exported by the gha or registry backends, so each hosted runner starts with an empty mount. The documented workaround is the buildkit-cache-dance action, which saves and restores the mount directory with actions/cache.

Do I need docker/setup-buildx-action for cache-to type=gha?

Yes, in practice. The gha and registry exporters need a builder that supports them. The default docker driver supports them only with the containerd image store enabled, so docker/setup-buildx-action creates a docker-container builder that works everywhere. Skipping the action usually produces an unsupported-driver error.

How big can the GitHub Actions Docker cache get?

GitHub’s default limit is 10 GB per repository, shared by every cache entry in that repo. Entries not accessed for over 7 days are removed. Once the limit is reached, GitHub saves the new cache and evicts the oldest-accessed entries. Repository admins can raise the limit, with extra storage billed.

Why does my Docker build ignore the cache even though cache-from is set?

The cache misses when something above the missed layer changed. Common causes are a COPY . . before dependency install, a build ARG that changes every run, no .dockerignore, a moved base image tag, or a different platform. Run --progress=plain and find the first layer without a CACHED line.

Can I use registry cache with GitLab CI or Forgejo Actions?

Yes. The registry backend is plain docker buildx build with --cache-from and --cache-to flags, so it runs on any CI that can reach a registry. You need a builder that can export cache (usually the docker-container driver) and a registry login. The gha backend works only inside GitHub Actions.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Previous Post
Btrfs Send/Receive: Incremental Backups
Next Post
Sync Game Saves Without a Cloud Account

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts