The Sync Is Red and It’s 2 AM
You pushed a one-line image bump before bed. Now the dashboard is red, the pager is loud, and the cluster refuses to get better on its own. You have two ways out, and picking the wrong one turns a five-minute fix into an hour.
git revert is the default rollback in GitOps, in both Argo CD and Flux. Git is the source of truth, so you fix the truth. The controller-side tools (argocd app rollback, Flux HelmRelease remediation) are the emergency brake. They work, and each has a sharp edge: Argo’s rollback refuses to run while auto-sync is on, and Flux has no generic rollback command at all. You suspend, revert, and reconcile.
I ran every failure in this post on a throwaway kind cluster with Argo CD v3.5.3 and Flux v2.9.5, so the error strings below are real output, not recall.
Full example: Clone the working files at github.com/KingPin/sumguy-examples/…/gitops-rollback-argocd-vs-flux
What “Broken” Actually Looks Like
A bad sync fails in one of four ways, and they look different in each tool.
- Bad image tag. The manifest is valid, the apply succeeds, and the new pod sits in
ImagePullBackOff. The rolling update keeps the old pods alive. - Broken manifest. The API server rejects the apply. Nothing changes in the cluster.
- Failing Helm upgrade. The release goes to
failedand the chart may be half applied. - Stuck operation. A sync retries forever, or a hook never finishes, and nothing else can sync until it stops.
The first one is the sneaky one. Argo CD reported this after I pushed nginx:1.29-alpine-broken and synced:
GROUP KIND NAMESPACE NAME STATUS HEALTH HOOK MESSAGEapps Deployment argo-demo web Synced Progressing deployment.apps/web configuredSynced. Progressing. A sync that “succeeded” while a pod is in ImagePullBackOff is like a forklift that delivered the couch to the wrong loading dock and got a signature anyway. The controller did its job. Your job is to notice.
Argo CD: Revert First, Rollback If You Must
The default: git revert
git revert --no-edit HEADgit pushargocd app get web --refreshWith auto-sync on, the controller picks up the revert on its next poll (every two to three minutes by default: timeout.reconciliation is 120s plus up to 60s of jitter), or immediately with --refresh, and applies it. In my test the app went from Progressing to Synced to main (73e2b9a) and Healthy in under 30 seconds after the refresh. The old pods never went away, so there was no downtime.
Revert is the right move because the repo now says what the cluster runs. Nothing to reconcile later, and your teammates see the fix in git log.
The emergency brake: argocd app rollback
When Git is slow, or the bad commit is buried under three good ones, use the sync history:
argocd app history webID DATE REVISION0 2026-09-29 15:56:16 -0400 EDT main (36bd52b)1 2026-09-29 15:57:25 -0400 EDT main (36bd52b)2 2026-09-29 15:58:48 -0400 EDT main (ce2c8b8)3 2026-09-29 15:59:29 -0400 EDT main (59b76dd)Entry 1 is a sync of an unchanged revision. Every sync gets an ID, even a no-op, so read the revision column, not the ID count. Argo keeps 10 entries by default (spec.revisionHistoryLimit: 10 in the Application reference manifest), which is also how far back you can go.
Roll back by history ID, not by revision:
argocd app rollback web 2The bad pods vanished and the app went Healthy. But look at what Argo says next:
Sync Status: OutOfSync from main (59b76dd)Health Status: HealthyThe cluster runs commit ce2c8b8. Git still says 59b76dd. Argo tells you the truth: the cluster and the repo disagree. A rollback does not change Git, so you now owe the repo a revert.
The trap: auto-sync
Turn on auto-sync and try the same rollback:
{"level":"fatal","msg":"rpc error: code = FailedPrecondition desc = rollback cannot be initiated when auto-sync is enabled"}Argo refuses. That is the polite failure. The rude one happens when you flip auto-sync on after a manual rollback. I did exactly that, and the controller saw OutOfSync from main (59b76dd) and immediately synced the bad commit again. A fresh ErrImagePull pod appeared five seconds later.
So the emergency brake needs a sequence:
argocd app set web --sync-policy noneargocd app rollback web 2# now fix Gitgit revert --no-edit <bad-sha> && git pushargocd app set web --sync-policy automatedRevert before you re-enable auto-sync. Skip that order and your 2 AM self will be reading the same dashboard at 2:07.
Broken manifests and stuck syncs
A typo like replicas: two gets past git push and dies at the API server. With syncPolicy.retry set, Argo does not give up on the first failure:
syncPolicy: automated: prune: true selfHeal: true retry: limit: 3 backoff: duration: 5s factor: 2 maxDuration: 1mThe operation stays in Running while it retries:
Running|one or more objects failed to apply, reason: error when patching "/dev/shm/495336992": "" is invalid: patch: Invalid value: "": unrecognized type: int32. Retrying attempt #3 at 8:05PM.The docs say limit: -1 means unlimited retries. That is a good way to build a sync that never ends. Set a limit. If you are already in the loop, stop it:
argocd app terminate-op webApplication 'web' operation terminatingA few seconds later the phase read Failed with Operation terminated (retried 3 times). One more fact from the auto-sync docs: Argo will not retry a failed automated sync against the same commit and parameters. Terminating the operation and then waiting for it to try again gets you nothing. Push a fix.
Flux: Suspend, Revert, Reconcile
Flux has no rollback command. That sounds like a gap until you notice that the workflow is the same three verbs every time.
The default: suspend, revert, reconcile
The Kustomization below has wait: true and timeout: 60s, so Flux runs health checks on everything it applies:
apiVersion: kustomize.toolkit.fluxcd.io/v1kind: Kustomizationmetadata: name: demo namespace: flux-systemspec: interval: 30s path: ./apps/flux-demo prune: true wait: true timeout: 60s sourceRef: kind: GitRepository name: homelabI pushed the same bad image tag. After the timeout, flux get kustomization demo said:
NAME REVISION SUSPENDED READY MESSAGEdemo main@sha1:b1b88314 False False health check failed after 1m0.007258489s: timeout waiting for: [Deployment/flux-demo/web status: 'InProgress']Without wait (or healthChecks) Flux would have reported success, exactly like Argo did. With it, you get a red Ready condition, which is the alert you want. Flux did not roll anything back. The old pod kept serving, and the controller kept retrying every interval.
Stop the retries while you fix things, then revert and reconcile:
flux suspend kustomization demogit revert --no-edit HEAD && git pushflux resume kustomization demoflux reconcile kustomization demo --with-source✔ kustomization resumed◎ waiting for Kustomization reconciliation✔ Kustomization demo reconciliation completed✔ applied revision main@sha1:3530e915fad42643a4364a66b45b306b2db77d65◎ waiting for GitRepository reconciliation✔ fetched revision main@sha1:3530e915fad42643a4364a66b45b306b2db77d65Per the Kustomization docs, a suspended Kustomization does not apply new source revisions and pauses drift correction. That is why you suspend first: it stops Flux from fighting you while you work, for example if you patch a live object by hand to stop the bleeding.
--with-source matters. Without it, flux reconcile kustomization re-applies whatever the last fetched artifact was. If the source has not polled since your revert, you just re-apply the bad commit. With --with-source, Flux fetches the GitRepository first.
Flux HelmRelease: Remediation Is the Brake
Helm gives Flux something a plain Kustomization lacks: a release history it can roll back to. The knobs live in spec.upgrade.remediation and spec.install.remediation.
apiVersion: helm.toolkit.fluxcd.io/v2kind: HelmReleasemetadata: name: demo namespace: flux-systemspec: interval: 30s releaseName: demo targetNamespace: flux-helm storageNamespace: flux-helm install: createNamespace: true remediation: retries: 1 timeout: 60s chart: spec: chart: ./charts/demo reconcileStrategy: Revision sourceRef: kind: GitRepository name: homelab values: image: tag: 1.29-alpine upgrade: remediation: retries: 2 strategy: rollback remediateLastFailure: true # defaults to true anyway when retries > 0Defaults from the HelmRelease docs, because they surprise people:
spec.timeoutdefaults to5m0s. That is how long a bad image tag hangs the upgrade before Flux calls it failed.upgrade.remediation.retriesdefaults to0, and a negative number means infinite retries.upgrade.remediation.strategydefaults torollback. The other value isuninstall, which reinstalls afterward.upgrade.remediation.remediateLastFailuredefaults tofalseunlessretriesis above0.install.remediationworks the same way, but the remediation is always an uninstall, andremediateLastFailuredefaults tofalse.
The retries default is the real gotcha. With the defaults (retries: 0), a failed upgrade is left as a failed release. Nothing rolls back. You have to opt in.
I bumped the tag to 1.29-alpine-broken with retries: 2. The controller ran the upgrade, hit the 60 second timeout, rolled back, tried again, rolled back, tried a third time, and rolled back the last failure because remediateLastFailure was on:
REVISION STATUS DESCRIPTION10 failed Upgrade "demo" failed: timeout waiting for: [Deployment/flux-helm/demo status: 'InProgress']11 superseded Rollback to 912 failed Upgrade "demo" failed: timeout waiting for: [Deployment/flux-helm/demo status: 'InProgress']13 superseded Rollback to 1114 failed Upgrade "demo" failed: timeout waiting for: [Deployment/flux-helm/demo status: 'InProgress']15 deployed Rollback to 13The final conditions:
Stalled=True RetriesExceeded: Failed to upgrade after 3 attempt(s)Remediated=True RollbackSucceeded: Helm rollback to previous release flux-helm/demo.v13 with chart [email protected]+4620db5a94fb succeededThree attempts for retries: 2: the first try plus two retries. The cluster is healthy again, which is the good news. The bad news is Stalled=True. The release is fine and Git is still wrong. The controller has given up. A new commit, a changed chart, or flux reconcile helmrelease demo --reset is what moves it again.
Fix Git anyway:
git revert --no-edit HEAD && git pushflux reconcile kustomization demo-helm --with-sourceflux reconcile helmrelease demodemo 0.1.0+d6e7d22aacc9 False True Helm upgrade succeeded for release flux-helm/demo.v18 with chart [email protected]+d6e7d22aacc9Remediation does not cover your plain manifests. My bad image tag in the Kustomization stayed broken until I reverted. It only helps when a Helm action fails.
Since Flux v2.7, HelmRelease also has upgrade.strategy.name (and install.strategy.name): RemediateOnFailure (the default) or RetryOnFailure. The second retries on strategy.retryInterval (default 5m) with no remediation. Keep RemediateOnFailure if you want the brake.
Argo CD vs Flux: The Side-by-Side
| Situation | Argo CD | Flux |
|---|---|---|
| Default fix | git revert, push, refresh | git revert, push, flux reconcile ... --with-source |
| Emergency brake | argocd app rollback <app> <ID> (auto-sync off) | HelmRelease upgrade.remediation (Helm only) |
| Pause the controller | argocd app set <app> --sync-policy none | flux suspend kustomization <name> |
| Bad image, valid YAML | Synced, Progressing | Ready=False only if wait or healthChecks is set |
| Stuck operation | argocd app terminate-op <app> | Wait out spec.timeout, or suspend |
| Retries | syncPolicy.retry with backoff | spec.retryInterval, remediation.retries |
Argo gives you a real rollback command and a sync history you can inspect. Flux gives you fewer moving parts. I lean toward Flux for pure GitOps discipline and Argo when I want to click a button at 2 AM and see a diff. The comparison post covers the rest of that choice.
The 2 AM Runbook
Copy this into your wiki. Your 2 AM self will thank you.
- Stop the bleeding. Argo:
argocd app set <app> --sync-policy none. Flux:flux suspend kustomization <name>(orflux suspend helmrelease <name>). - Kill any stuck operation. Argo:
argocd app terminate-op <app>. - Check what is actually running.
kubectl get pods -n <ns>andargocd app get <app>orflux get all. A rolling update usually keeps the old pods alive, so you may have more time than the dashboard implies. - Restore service if it is down. Argo:
argocd app history <app>, thenargocd app rollback <app> <ID>. Flux has no rollback command, so do step 5 fast. A revert plusflux reconcile ... --with-sourceis the rollback. - Fix Git.
git revert --no-edit <bad-sha>and push. Do not hand-edit the manifest under pressure. - Resume and verify. Argo:
argocd app set <app> --sync-policy automated, thenargocd app get <app> --refresh. Flux:flux resume kustomization <name>, thenflux reconcile kustomization <name> --with-source. - Confirm Git and cluster agree. Argo shows
Synced to main (<sha>). Flux showsApplied revision: main@sha1:<sha>.
The order in step 6 is not optional. Re-enable auto-sync while the repo still holds the bad commit and Argo redeploys it.
Common Questions
Can I use argocd app rollback with auto-sync enabled?
No. Argo CD v3.5.3 rejects the command with rollback cannot be initiated when auto-sync is enabled. Disable auto-sync first with argocd app set <app> --sync-policy none, roll back, then revert the commit in Git before you turn auto-sync back on, or the controller redeploys the bad commit.
Does a Flux HelmRelease roll back a bad Kustomization change?
No. HelmRelease remediation only reacts to failed Helm install and upgrade actions. A bad image tag in a plain Kustomization keeps failing its health check until you revert the commit. Flux reports Ready=False for that Kustomization but leaves the previous pods running, so you get an alert and no automatic fix.
How many rollback points does Argo CD keep?
Argo CD keeps 10 history entries per Application by default, set by spec.revisionHistoryLimit (the Application reference manifest shows 10). Every sync adds an entry, including syncs of an unchanged revision, so a busy Application can push a good revision out of history.
Why did my HelmRelease not roll back after a failed upgrade?
The default is upgrade.remediation.retries: 0, and remediateLastFailure defaults to false unless retries is above 0. With both defaults, Flux leaves the failed release in place. Set retries to at least 1 and keep strategy: rollback to get an automatic rollback after a failed upgrade.