Skip to content
Go back

GitOps Rollback: Argo CD vs Flux

By KingPin 12 min read
GitOps Rollback: Argo CD vs Flux
Contents

The Sync Is Red and It’s 2 AM

You pushed a one-line image bump before bed. Now the dashboard is red, the pager is loud, and the cluster refuses to get better on its own. You have two ways out, and picking the wrong one turns a five-minute fix into an hour.

git revert is the default rollback in GitOps, in both Argo CD and Flux. Git is the source of truth, so you fix the truth. The controller-side tools (argocd app rollback, Flux HelmRelease remediation) are the emergency brake. They work, and each has a sharp edge: Argo’s rollback refuses to run while auto-sync is on, and Flux has no generic rollback command at all. You suspend, revert, and reconcile.

I ran every failure in this post on a throwaway kind cluster with Argo CD v3.5.3 and Flux v2.9.5, so the error strings below are real output, not recall.

Full example: Clone the working files at github.com/KingPin/sumguy-examples/…/gitops-rollback-argocd-vs-flux

What “Broken” Actually Looks Like

A bad sync fails in one of four ways, and they look different in each tool.

  1. Bad image tag. The manifest is valid, the apply succeeds, and the new pod sits in ImagePullBackOff. The rolling update keeps the old pods alive.
  2. Broken manifest. The API server rejects the apply. Nothing changes in the cluster.
  3. Failing Helm upgrade. The release goes to failed and the chart may be half applied.
  4. Stuck operation. A sync retries forever, or a hook never finishes, and nothing else can sync until it stops.

The first one is the sneaky one. Argo CD reported this after I pushed nginx:1.29-alpine-broken and synced:

GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGE
apps Deployment argo-demo web Synced Progressing deployment.apps/web configured

Synced. Progressing. A sync that “succeeded” while a pod is in ImagePullBackOff is like a forklift that delivered the couch to the wrong loading dock and got a signature anyway. The controller did its job. Your job is to notice.

Argo CD: Revert First, Rollback If You Must

The default: git revert

Terminal window
git revert --no-edit HEAD
git push
argocd app get web --refresh

With auto-sync on, the controller picks up the revert on its next poll (every two to three minutes by default: timeout.reconciliation is 120s plus up to 60s of jitter), or immediately with --refresh, and applies it. In my test the app went from Progressing to Synced to main (73e2b9a) and Healthy in under 30 seconds after the refresh. The old pods never went away, so there was no downtime.

Revert is the right move because the repo now says what the cluster runs. Nothing to reconcile later, and your teammates see the fix in git log.

The emergency brake: argocd app rollback

When Git is slow, or the bad commit is buried under three good ones, use the sync history:

Terminal window
argocd app history web
ID DATE REVISION
0 2026-09-29 15:56:16 -0400 EDT main (36bd52b)
1 2026-09-29 15:57:25 -0400 EDT main (36bd52b)
2 2026-09-29 15:58:48 -0400 EDT main (ce2c8b8)
3 2026-09-29 15:59:29 -0400 EDT main (59b76dd)

Entry 1 is a sync of an unchanged revision. Every sync gets an ID, even a no-op, so read the revision column, not the ID count. Argo keeps 10 entries by default (spec.revisionHistoryLimit: 10 in the Application reference manifest), which is also how far back you can go.

Roll back by history ID, not by revision:

Terminal window
argocd app rollback web 2

The bad pods vanished and the app went Healthy. But look at what Argo says next:

Sync Status: OutOfSync from main (59b76dd)
Health Status: Healthy

The cluster runs commit ce2c8b8. Git still says 59b76dd. Argo tells you the truth: the cluster and the repo disagree. A rollback does not change Git, so you now owe the repo a revert.

The trap: auto-sync

Turn on auto-sync and try the same rollback:

{"level":"fatal","msg":"rpc error: code = FailedPrecondition desc = rollback cannot be initiated when auto-sync is enabled"}

Argo refuses. That is the polite failure. The rude one happens when you flip auto-sync on after a manual rollback. I did exactly that, and the controller saw OutOfSync from main (59b76dd) and immediately synced the bad commit again. A fresh ErrImagePull pod appeared five seconds later.

So the emergency brake needs a sequence:

Terminal window
argocd app set web --sync-policy none
argocd app rollback web 2
# now fix Git
git revert --no-edit <bad-sha> && git push
argocd app set web --sync-policy automated

Revert before you re-enable auto-sync. Skip that order and your 2 AM self will be reading the same dashboard at 2:07.

Broken manifests and stuck syncs

A typo like replicas: two gets past git push and dies at the API server. With syncPolicy.retry set, Argo does not give up on the first failure:

application.yaml (excerpt)
syncPolicy:
automated:
prune: true
selfHeal: true
retry:
limit: 3
backoff:
duration: 5s
factor: 2
maxDuration: 1m

The operation stays in Running while it retries:

Running|one or more objects failed to apply, reason: error when patching "/dev/shm/495336992": "" is invalid: patch: Invalid value: "": unrecognized type: int32. Retrying attempt #3 at 8:05PM.

The docs say limit: -1 means unlimited retries. That is a good way to build a sync that never ends. Set a limit. If you are already in the loop, stop it:

Terminal window
argocd app terminate-op web
Application 'web' operation terminating

A few seconds later the phase read Failed with Operation terminated (retried 3 times). One more fact from the auto-sync docs: Argo will not retry a failed automated sync against the same commit and parameters. Terminating the operation and then waiting for it to try again gets you nothing. Push a fix.

Flux: Suspend, Revert, Reconcile

Flux has no rollback command. That sounds like a gap until you notice that the workflow is the same three verbs every time.

The default: suspend, revert, reconcile

The Kustomization below has wait: true and timeout: 60s, so Flux runs health checks on everything it applies:

kustomizations.yaml (excerpt)
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: demo
namespace: flux-system
spec:
interval: 30s
path: ./apps/flux-demo
prune: true
wait: true
timeout: 60s
sourceRef:
kind: GitRepository
name: homelab

I pushed the same bad image tag. After the timeout, flux get kustomization demo said:

NAME REVISION SUSPENDED READY MESSAGE
demo main@sha1:b1b88314 False False health check failed after 1m0.007258489s: timeout waiting for: [Deployment/flux-demo/web status: 'InProgress']

Without wait (or healthChecks) Flux would have reported success, exactly like Argo did. With it, you get a red Ready condition, which is the alert you want. Flux did not roll anything back. The old pod kept serving, and the controller kept retrying every interval.

Stop the retries while you fix things, then revert and reconcile:

Terminal window
flux suspend kustomization demo
git revert --no-edit HEAD && git push
flux resume kustomization demo
flux reconcile kustomization demo --with-source
✔ kustomization resumed
◎ waiting for Kustomization reconciliation
✔ Kustomization demo reconciliation completed
✔ applied revision main@sha1:3530e915fad42643a4364a66b45b306b2db77d65
◎ waiting for GitRepository reconciliation
✔ fetched revision main@sha1:3530e915fad42643a4364a66b45b306b2db77d65

Per the Kustomization docs, a suspended Kustomization does not apply new source revisions and pauses drift correction. That is why you suspend first: it stops Flux from fighting you while you work, for example if you patch a live object by hand to stop the bleeding.

--with-source matters. Without it, flux reconcile kustomization re-applies whatever the last fetched artifact was. If the source has not polled since your revert, you just re-apply the bad commit. With --with-source, Flux fetches the GitRepository first.

Flux HelmRelease: Remediation Is the Brake

Helm gives Flux something a plain Kustomization lacks: a release history it can roll back to. The knobs live in spec.upgrade.remediation and spec.install.remediation.

helmrelease.yaml
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: demo
namespace: flux-system
spec:
interval: 30s
releaseName: demo
targetNamespace: flux-helm
storageNamespace: flux-helm
install:
createNamespace: true
remediation:
retries: 1
timeout: 60s
chart:
spec:
chart: ./charts/demo
reconcileStrategy: Revision
sourceRef:
kind: GitRepository
name: homelab
values:
image:
tag: 1.29-alpine
upgrade:
remediation:
retries: 2
strategy: rollback
remediateLastFailure: true # defaults to true anyway when retries > 0

Defaults from the HelmRelease docs, because they surprise people:

The retries default is the real gotcha. With the defaults (retries: 0), a failed upgrade is left as a failed release. Nothing rolls back. You have to opt in.

I bumped the tag to 1.29-alpine-broken with retries: 2. The controller ran the upgrade, hit the 60 second timeout, rolled back, tried again, rolled back, tried a third time, and rolled back the last failure because remediateLastFailure was on:

REVISION STATUS DESCRIPTION
10 failed Upgrade "demo" failed: timeout waiting for: [Deployment/flux-helm/demo status: 'InProgress']
11 superseded Rollback to 9
12 failed Upgrade "demo" failed: timeout waiting for: [Deployment/flux-helm/demo status: 'InProgress']
13 superseded Rollback to 11
14 failed Upgrade "demo" failed: timeout waiting for: [Deployment/flux-helm/demo status: 'InProgress']
15 deployed Rollback to 13

The final conditions:

Stalled=True RetriesExceeded: Failed to upgrade after 3 attempt(s)
Remediated=True RollbackSucceeded: Helm rollback to previous release flux-helm/demo.v13 with chart [email protected]+4620db5a94fb succeeded

Three attempts for retries: 2: the first try plus two retries. The cluster is healthy again, which is the good news. The bad news is Stalled=True. The release is fine and Git is still wrong. The controller has given up. A new commit, a changed chart, or flux reconcile helmrelease demo --reset is what moves it again.

Fix Git anyway:

Terminal window
git revert --no-edit HEAD && git push
flux reconcile kustomization demo-helm --with-source
flux reconcile helmrelease demo
demo 0.1.0+d6e7d22aacc9 False True Helm upgrade succeeded for release flux-helm/demo.v18 with chart [email protected]+d6e7d22aacc9

Remediation does not cover your plain manifests. My bad image tag in the Kustomization stayed broken until I reverted. It only helps when a Helm action fails.

Since Flux v2.7, HelmRelease also has upgrade.strategy.name (and install.strategy.name): RemediateOnFailure (the default) or RetryOnFailure. The second retries on strategy.retryInterval (default 5m) with no remediation. Keep RemediateOnFailure if you want the brake.

Argo CD vs Flux: The Side-by-Side

SituationArgo CDFlux
Default fixgit revert, push, refreshgit revert, push, flux reconcile ... --with-source
Emergency brakeargocd app rollback <app> <ID> (auto-sync off)HelmRelease upgrade.remediation (Helm only)
Pause the controllerargocd app set <app> --sync-policy noneflux suspend kustomization <name>
Bad image, valid YAMLSynced, ProgressingReady=False only if wait or healthChecks is set
Stuck operationargocd app terminate-op <app>Wait out spec.timeout, or suspend
RetriessyncPolicy.retry with backoffspec.retryInterval, remediation.retries

Argo gives you a real rollback command and a sync history you can inspect. Flux gives you fewer moving parts. I lean toward Flux for pure GitOps discipline and Argo when I want to click a button at 2 AM and see a diff. The comparison post covers the rest of that choice.

The 2 AM Runbook

Copy this into your wiki. Your 2 AM self will thank you.

  1. Stop the bleeding. Argo: argocd app set <app> --sync-policy none. Flux: flux suspend kustomization <name> (or flux suspend helmrelease <name>).
  2. Kill any stuck operation. Argo: argocd app terminate-op <app>.
  3. Check what is actually running. kubectl get pods -n <ns> and argocd app get <app> or flux get all. A rolling update usually keeps the old pods alive, so you may have more time than the dashboard implies.
  4. Restore service if it is down. Argo: argocd app history <app>, then argocd app rollback <app> <ID>. Flux has no rollback command, so do step 5 fast. A revert plus flux reconcile ... --with-source is the rollback.
  5. Fix Git. git revert --no-edit <bad-sha> and push. Do not hand-edit the manifest under pressure.
  6. Resume and verify. Argo: argocd app set <app> --sync-policy automated, then argocd app get <app> --refresh. Flux: flux resume kustomization <name>, then flux reconcile kustomization <name> --with-source.
  7. Confirm Git and cluster agree. Argo shows Synced to main (<sha>). Flux shows Applied revision: main@sha1:<sha>.

The order in step 6 is not optional. Re-enable auto-sync while the repo still holds the bad commit and Argo redeploys it.

Common Questions

Can I use argocd app rollback with auto-sync enabled?

No. Argo CD v3.5.3 rejects the command with rollback cannot be initiated when auto-sync is enabled. Disable auto-sync first with argocd app set <app> --sync-policy none, roll back, then revert the commit in Git before you turn auto-sync back on, or the controller redeploys the bad commit.

Does a Flux HelmRelease roll back a bad Kustomization change?

No. HelmRelease remediation only reacts to failed Helm install and upgrade actions. A bad image tag in a plain Kustomization keeps failing its health check until you revert the commit. Flux reports Ready=False for that Kustomization but leaves the previous pods running, so you get an alert and no automatic fix.

How many rollback points does Argo CD keep?

Argo CD keeps 10 history entries per Application by default, set by spec.revisionHistoryLimit (the Application reference manifest shows 10). Every sync adds an entry, including syncs of an unchanged revision, so a busy Application can push a good revision out of history.

Why did my HelmRelease not roll back after a failed upgrade?

The default is upgrade.remediation.retries: 0, and remediateLastFailure defaults to false unless retries is above 0. With both defaults, Flux leaves the failed release in place. Set retries to at least 1 and keep strategy: rollback to get an automatic rollback after a failed upgrade.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Next Post
Temporal vs Cron: Durable Home Lab Jobs

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts