Argo CD 3.x Upgrade Guide: What Changed vs v2 for Production

Argo CD 3.x Upgrade Guide: What Changed vs v2 for Production

Argo CD 3 Upgrade: What Changed vs v2 and a Safe Rollout Plan

If you still run Argo CD 2.x in production, you are running software the project no longer patches. Argo CD ships four minor releases a year and only the three most recent minors receive fixes; with 3.5 generally available since August 2026, that window is 3.5, 3.4 and 3.3, and the last 2.x line (2.14) left support when 3.2 shipped. The catch is that the jump from 2.14 to 3.0 is the one upgrade in the project’s history that deliberately changed defaults: RBAC on pod logs, how resources are tracked, which objects the controller watches, and how Dex maps identities. A careless Argo CD 3 upgrade does not fail loudly. It quietly locks engineers out of logs, flips applications to OutOfSync, or stops a sync that used to work.

This guide is written for platform teams who operate Argo CD for real fleets. It walks through every default that changed between v2.14 and 3.0, the breaking changes that landed in 3.1 through 3.5, the sharding and performance implications, and a staged, rollback-safe plan you can execute in a change window.

What this covers: the 2.14-to-3.0 breaking changes and their exact config keys, what each later minor adds to your checklist, how to pre-flight and validate an upgrade, how sharding interacts with it, and the failure modes to watch for.

Context and Background

Argo CD 3.0 was the first major version since 2021. The project’s own release announcement framed it as a clean-up: remove deprecated behavior, and turn hard-won operational advice into defaults while leaving an escape hatch for teams that want the old behavior. GA landed on 6 May 2025 according to the project’s release calendar, with 3.1 in August 2025, 3.2 in November 2025, and the cadence continuing on the first Tuesday of February, May, August and November. That calendar is why the support arithmetic matters: a cluster left on 2.14 has been out of the patch window for almost a year by now.

Semantic versioning tells you what to expect from each hop. The upgrade overview states that patch releases carry no breaking changes, minor releases may need workarounds documented in their own guide, and major releases contain backward-incompatible changes and warrant a backup of settings using the disaster recovery tooling. Practically, that means the 2.14 to 3.0 page is mandatory reading, and each minor page after it is a shorter but still real checklist. We covered the strategic choice between Argo CD and its main alternative in the Argo CD vs Flux GitOps decision record; this article assumes you have already chosen Argo CD and need to keep it current.

Two facts shape how you plan. First, Argo CD is a control plane that sits on your deploy path, so a botched upgrade is an availability event for every team that ships through it, though workloads keep running because Argo CD does not sit in the data path. Second, the defaults that changed in 3.0 are mostly security and scale hardening, which means rolling back the settings to v2 values is a legitimate transitional tactic, provided you record it as debt. The authoritative source for everything below is the official upgrade documentation; I cite config keys from it verbatim and label where I am inferring.

If you operate Argo CD at the edge or across industrial sites, the same upgrade discipline applies with extra constraints about intermittent connectivity, which we treated in the Argo CD and Flux GitOps for industrial fleets tutorial.

The Upgrade Map: From 2.14 to 3.5

The most useful mental model is a staircase, not a single jump. The project’s guidance is to read every intermediate upgrade page, and the safest supported path from an older 2.x release is: land on 2.14 first, take the major step to 3.0 (or straight to a supported 3.x patch), then move up one minor at a time in non-production before production.

Direct answer: the biggest Argo CD v3 breaking changes are all in 3.0: fine-grained RBAC no longer inherits, log access is always RBAC-checked, annotation tracking is the default, new resource exclusions hide noisy kinds, three legacy metrics are gone, Dex identities changed, and repositories can no longer live in argocd-cm. Later minors add narrower changes such as server-side apply for the ApplicationSet CRD in 3.3.

Argo CD 3 upgrade path from 2.14 to 3.5 showing the breaking change gates at each minor

Figure 1: The upgrade staircase. Each gate lists the one change most likely to bite a production install.

The figure above compresses roughly twenty documented changes into the handful that most often cause an incident. It is deliberately opinionated, so let me justify the ordering.

Why 3.0 is the only real cliff

Between 3.0 and 3.5 the upgrade pages list changes that are additive, deprecations, or narrow (a health-check behavior, a CRD size limit, a cluster version string). Only 3.0 changes what existing, working configuration does without you touching it. That is the property that makes it dangerous: there is no error at upgrade time, only different behavior. The right response is to convert every silent change into an explicit decision before the upgrade, which is what the pre-flight section below does.

Pin a version, not a channel

A frequent anti-pattern is installing from a stable manifest URL or an unpinned Helm chart range. For a control plane, pin the exact image tag and the full manifest, and apply the whole manifest rather than only bumping the image. The upgrade overview is explicit that manifest changes can include critical parameter modifications, and that the documented apply commands use server-side apply with force-conflicts because the CRDs are too large for client-side apply. That last point is not a 3.3 novelty for ApplicationSet alone; treat server-side apply as the default mechanism for all Argo CD CRDs.

Patch-level hygiene

Within a minor, you should track the latest patch. As of this writing the project’s releases page shows v3.5.2 and v3.4.8 published on 27 August 2026, and v3.3.14 on 12 August 2026. Landing on the newest patch of your target minor, rather than the .0, avoids a class of first-release bugs. This is a general release-engineering habit rather than a documented Argo CD rule, so treat it as my recommendation.

What Changed in Argo CD 3.0: The Nine Breaking Changes and Eight Default Shifts

The 2.14-to-3.0 guide separates breaking changes from default changes. I keep that split because breaking changes need a code or config edit, while default changes need only a decision. Everything here is from the official guide unless marked otherwise.

1. Fine-grained RBAC no longer inherits

In v2, an update or delete grant on an application implicitly covered the resources the application manages. In 3.0, update and delete actions no longer inherit down to sub-resources. To grant them you write explicit policies using the <action>/<group>/<kind>/<namespace>/<name> form, for example a role that may delete only Pods in one application:

p, role:pod-janitor, applications, delete/*/Pod/*/*, default/prod-app, allow

To restore the old inheritance you set server.rbac.disableApplicationFineGrainedRBACInheritance: "false" in argocd-cm. The RBAC documentation adds a subtlety that trips people: glob matching does not treat / as a separator, so a pattern needs all its path components or it may match more than you intend.

2. Log access is always enforced

Previously, pod-log access was RBAC-checked only if you set server.rbac.log.enforce.enable. In 3.0 that flag is removed and enforcement is unconditional. Anyone whose role lacks a logs, get grant loses log access in the UI and CLI. The quick fix is p, role:<YOUR_ROLE>, logs, get, */*, allow; the recommended fix is to grant it per role and per project, because */* gives every application’s logs to that role.

3. New default resource exclusions

resource.exclusions in argocd-cm now ships with defaults that hide high-churn kinds: Endpoints, EndpointSlice, Lease, certificate request objects, policy reports, and objects from Cilium and Kyverno. The effect is fewer watch events and less memory in the controller. The risk is that if you deliberately manage one of these kinds through Argo CD, it disappears from the resource tree. Override or delete the setting to preserve v2 behavior.

4. Legacy metrics removed

argocd_app_sync_status, argocd_app_health_status and argocd_app_created_time are gone. If ARGOCD_LEGACY_CONTROLLER_METRICS was ever enabled in your environment, your dashboards and alerts probably use them. Rewrite queries against argocd_app_info, which exposes sync and health status as labels. This is the change most likely to make a monitoring stack go silent, and a silent alert is worse than a failing one.

5. Dex subject claims

Dex-based SSO now identifies users by federated_claims.user_id rather than the sub claim. Any RBAC policy written against the old subject value stops matching. The guide suggests decoding existing sub values (they are base64-encoded structures) to find the underlying user ID and rewriting policies. If you use group claims exclusively, you may be unaffected, but verify it; do not assume.

6. Repositories out of argocd-cm

Repository and credential definitions can no longer live in argocd-cm under repositories, repository.credentials or helm.repositories. Convert them to labeled Secrets. The guide provides a kubectl jsonpath one-liner to detect leftovers, and I recommend running it in CI against your GitOps repository so it cannot regress.

7. ApplicationSet nested selectors

The applyNestedSelectors field is ignored and nested selectors are always applied. If you had it unset or false and relied on nested selectors being skipped, generators that combine (matrix, merge) may now produce fewer Applications. The guide points you to logs containing “ignoring nested selector” to find affected sets. Fewer generated Applications can mean pruning, which is the scary part; see the gotchas section.

8. Helm 3.17.1 semantics

The bundled Helm changed how null values in values.yaml behave: null now overrides objects rather than emitting a warning. Charts that used nulls to blank out subchart defaults can render different manifests. Diff your rendered output before and after (a recipe follows).

9. Annotation tracking by default

The most consequential default. Argo CD 3 tracks ownership with the argocd.argoproj.io/tracking-id annotation, formatted as <app>:<group>/<kind>:<namespace>/<name>, instead of the app.kubernetes.io/instance label. It exists because labels are limited to 63 characters and collide with other tools that also set that label. Set application.resourceTrackingMethod: label to keep v2 behavior; valid values are label, annotation+label and annotation. Applications using ApplyOutOfSyncOnly=true need an explicit sync after upgrade so annotations get written.

The default shifts that need a decision, not an edit

The guide lists further defaults that change behavior without breaking configs:

  • Health no longer persisted in the Application CR (controller.resource.health.persist, set true to restore). This reduces write load on the Application objects and etcd. Tools that read health straight out of the Application status will see a new resourceHealthSource field pattern.
  • Status field ignored for all resources in diffs, not only CRDs (resource.compareoptions with ignoreResourceStatusField; set crd for v2).
  • ignoreDifferences applied during updates (ignoreDifferencesOnResourceUpdates; set false for v2).
  • spec.preserveUnknownFields no longer ignored by default on CRDs, so add an ignoreDifferences entry if CRDs flip OutOfSync.
  • Empty plugin environment variables are now passed to config management plugins rather than omitted.
  • In-cluster can be disabled with cluster.inClusterEnabled: "false"; apps targeting it go to Unknown.
  • High-churn update filtering was added to resource.customizations.ignoreResourceUpdates, reducing reconcile load.
  • Project API sanitization removed project-scoped repository and cluster credentials from API responses.

That is the whole 3.0 surface. Before applying anything, it is worth seeing how these interact, starting with RBAC, where two separate changes compound.

Deeper Analysis: Making the Silent Changes Loud

The goal of a good GitOps upgrade plan is to turn every implicit behavior into an observed fact before production sees it. Three areas deserve a walk-through: RBAC, resource tracking, and the rendered-manifest diff.

RBAC: two changes that compound

Consider a typical v2 setup. A developer role has applications, update, team-a/*, allow and applications, get, team-a/*, allow. In v2 that update grant covered sub-resource actions such as restarting a Deployment or deleting a Pod through the UI. There was no logs grant because log enforcement was off. After the upgrade to 3.0, that developer can still sync the application, but Pod deletion and resource actions now need update/* or delete/* policies, and log tabs return permission errors.

Argo CD RBAC evaluation before and after the 3.0 upgrade for application update, sub-resource delete and logs

Figure 2: How one request is authorized in v2 versus 3.x. Sub-resource actions and logs each need their own explicit allow.

The evaluation logic in the figure is the part to internalize. In 3.x, a request for a sub-resource action is a distinct check against <action>/<group>/<kind>/<namespace>/<name>, and a log read is a distinct check against the logs resource. Neither is satisfied by the application-level grant.

A migration policy fragment that reproduces v2 for a developer role while remaining reviewable looks like this:

# argocd-rbac-cm policy.csv (excerpt)
p, role:developer, applications, get,            team-a/*, allow
p, role:developer, applications, sync,           team-a/*, allow
p, role:developer, applications, update/*,       team-a/*, allow
p, role:developer, applications, delete/*/Pod/*/*, team-a/*, allow
p, role:developer, logs,         get,            team-a/*, allow
g, team-a-devs, role:developer

Validate before you deploy it. The argocd admin settings rbac can command evaluates a policy file offline, so you can assert that a developer can read logs and cannot delete Deployments, and run those assertions in CI:

argocd admin settings rbac can role:developer get logs 'team-a/checkout' \
  --policy-file policy.csv
argocd admin settings rbac can role:developer delete 'applications' 'team-a/checkout' \
  --policy-file policy.csv   # expect No

One trap: policy.default. The documentation states that all authenticated users receive at least the permissions of the default policy and that a deny rule cannot remove it. If you set policy.default: role:readonly to be friendly, note that in 3.x whether readonly can see logs depends on the built-in policy; test it rather than assume.

The escape hatch, server.rbac.disableApplicationFineGrainedRBACInheritance: "false", is acceptable for the first upgrade window. My recommendation is to time-box it: set an expiry ticket, and remove it once explicit update/* policies are in place. Leaving the v2 inheritance on indefinitely defeats the security reason the change was made, which is that a role able to update an application should not automatically be able to delete every object it owns.

Resource tracking: the label-to-annotation migration

Under label tracking, Argo CD stamps app.kubernetes.io/instance: <app-name> on every object and finds its resources by that label. Under annotation tracking it writes argocd.argoproj.io/tracking-id: <app>:<group>/<kind>:<namespace>/<name>, which encodes the identity of the object and so avoids two failure classes: labels truncating at 63 characters and other tools such as Helm or Kustomize commonly setting the same instance label.

The migration is safe in principle because Argo CD adds the annotation on the next sync, but three details matter.

First, until an application syncs, its resources have no tracking annotation. The controller still recognizes them, but a resource that exists in the cluster with a label from an old owner can be shown as belonging to two apps. Sync each application once after the upgrade. Applications using ApplyOutOfSyncOnly=true will not rewrite unchanged resources, so they require an explicit full sync; the guide calls this out.

Second, clusters with multiple Argo CD installations must set installationID so each instance writes its own argocd.argoproj.io/installation-id annotation. Without it, two instances tracking the same namespace can fight over ownership of resources with identical names.

Third, if you cannot sync everything at once, use annotation+label. The app.kubernetes.io/instance label is then written for informational use and external tools, while ownership is decided by the annotation. This is my preferred stepping stone for estates that have other tools reading that label, such as cost-allocation or service-mesh dashboards.

Sequence for switching Argo CD resource tracking during the 3 upgrade with a checkpoint and rollback

Figure 3: The tracking cutover as a sequence, with the checkpoint that makes rollback a one-line revert.

The sequence in the figure is intentionally boring: back up, pin the old behavior, upgrade, verify, then change one setting and sync. Separating “upgrade the binaries” from “change the tracking method” means that if something is wrong you know which change caused it.

Diff the rendered manifests, not just the config

Helm 3.17.1 semantics, Kustomize bumps, resource exclusion defaults and ignoreDifferences behavior all affect what Argo CD thinks the desired and live states are. The most reliable pre-flight is to render every application with the old and new repo-server version and diff the output. The CLI can render locally, and for a large estate a loop works:

# render each app with the currently installed server, save output
for app in $(argocd app list -o name); do
  argocd app manifests "$app" --source live   > "before/${app//\//_}.live.yaml"
  argocd app manifests "$app" --source git    > "before/${app//\//_}.git.yaml"
done
# after upgrading a staging instance pointed at the same repos, repeat into after/
diff -ru before after | less

Anything that changes in the git render is a genuine renderer difference (typically Helm null handling). Anything that changes only in the live comparison is a diff-customization default. Both categories are decisions, but they have different owners: chart authors for the first, platform owners for the second.

Later minors: what 3.1 through 3.5 add

After the cliff, each minor is a short checklist. Grouping the documented changes by risk:

3.1 (August 2025). The project API response is sanitized to remove credentials of project-scoped repositories and clusters (GHSA-786q-9hcg-v9ff). Static assets are protected against out-of-bounds symlinks, and external symlinks produce HTTP 500 errors. The v1 resource actions API is deprecated in favor of /api/v1/applications/{name}/resource/actions/v2 with a JSON body. OIDC login now expects an /auth/callback redirect URI registered on the identity provider. Helm moved to 3.18.4 and Kustomize to 5.7.0, and 17 additional resource types gained built-in health checks.

3.2 (November 2025). The source hydrator requires a non-root path per application, so repository root paths (empty or .) are no longer supported. A CronJob now contributes to health and can push an application to Degraded; a suspended CronJob stays healthy unless you customize resource.customizations.health.batch_CronJob, and the annotation argocd.argoproj.io/ignore-healthcheck: "true" excludes it. ApplicationSet status.resources is capped at 5,000 elements by default via applicationsetcontroller.status.max.resources.count. kustomize.version can be set in .argocd-source.yaml.

3.3 (February 2026). The important one: the ApplicationSet CRD now exceeds the client-side apply limit, so you must use server-side apply. If Argo CD manages itself, add ServerSideApply=true to sync options; for manual installs use kubectl apply -n argocd --server-side --force-conflicts. Helm-based upgrades are unaffected. The source hydrator moved to Git Notes for tracking hydration state and no longer cleans application paths, so stale files remain unless you remove them. ARGOCD_K8S_SERVER_SIDE_TIMEOUT now controls the Kubernetes server-side timeout while ARGOCD_K8S_TCP_TIMEOUT handles only TCP. Anonymous Settings API responses no longer expose resourceOverrides.

3.4 (May 2026). Cluster versions are stored as vMajor.Minor.Patch instead of Major.Minor, which matters for ApplicationSet cluster generators that reference argocd.argoproj.io/kubernetes-version. Application health shows Missing only when all resources are missing. GRPC_ENABLE_TXT_SERVICE_CONFIG defaults to false to avoid excess DNS lookups on dual-stack networks. Dex moved to 2.45.0 with ContinueOnConnectorFailure on by default, and the OpenTelemetry bump to 1.42.0 can rename metrics your Grafana dashboards depend on.

3.5 (August 2026). UI application names get a user-app- CSS class prefix; Helm 4.2.x makes OCI handling stricter (plain-HTTP registries need --insecure-oci-force-http); UI extensions must externalize react/jsx-runtime for React 19; gRPC event listing returns an Argo CD-defined EventList; and impersonation now applies to all server operations. New opt-in mTLS between components and the Source Integrity framework for Git signature verification headline the release, while Impersonation and the Source Hydrator graduated to beta.

If you are jumping from 2.14 straight to 3.4 or 3.5, you read all of these, but you only apply them once. That is an argument for not stopping at 3.0 unless you have a reason to: a fleet parked on 3.0 or 3.1 is again out of the patch window.

For clusters that use progressive delivery on top of Argo CD, remember that Argo Rollouts is versioned separately; our Argo Rollouts progressive delivery guide for edge fleets covers that component on its own, and the underlying cluster upgrade in Kubernetes 1.37 is a separate change you should not bundle into the same window.

Sharding and Scale: What the Upgrade Does to Controller Load

Most 3.0 defaults are scale features in disguise. Persisting health outside the Application CR, ignoring status in diffs, excluding Endpoints, EndpointSlice and Lease objects, and filtering high-churn updates all cut the number of events the application controller must process and the number of writes it makes to the Kubernetes API. I am not aware of a published, independent benchmark that quantifies the combined effect for a given fleet size, so I will not invent a percentage; measure your own before and after with the metrics described below.

Argo CD sharding is the other half of scale, and it is unchanged in spirit by the 3.x line, which is why it is worth checking during the upgrade window rather than after.

How controller sharding works

The application controller is a StatefulSet. Each replica owns a subset of clusters (not applications), and reconciles all applications targeting those clusters. You enable sharding by raising the replica count and setting ARGOCD_CONTROLLER_REPLICAS to the same number so each pod knows the total. The algorithm is selected with controller.sharding.algorithm in argocd-cmd-params-cm, the --sharding-method flag, or ARGOCD_CONTROLLER_SHARDING_ALGORITHM:

  • legacy: hashes the cluster UID. Distribution is not uniform, and it is the default.
  • round-robin: even distribution across shards. The docs mark it experimental and note that removing a cluster can reshuffle assignments.
  • consistent-hashing: consistent hashing with bounded loads. Also marked experimental; it reduces reshuffling when clusters or replicas change.

You can also pin a cluster to a shard by setting a shard field on its cluster Secret, independent of the algorithm. That is useful for isolating one very large or very noisy cluster.

Argo CD sharding topology with controller replicas, repo-server, Redis and cluster secrets after the 3 upgrade

Figure 4: Sharded application controllers own clusters; repo-servers and Redis are shared. Restarts during upgrade re-run shard assignment.

Two upgrade-time consequences follow. First, a StatefulSet rolls pods one at a time, and a restarting replica stops reconciling its clusters until it is back. With legacy hashing, an unlucky replica may own most of your large clusters, so a rolling restart can produce a longer blind spot than the replica count suggests. Second, if you change the algorithm and the replica count in the same window as the version upgrade, you will not be able to attribute a regression. Change one variable per window.

Queue tuning and the hydrator

Each replica processes reconciliation and sync through separate work queues. The documented flags and defaults are --status-processors (20), --operation-processors (10) and, when the source hydrator is enabled, --hydration-processors (5). Raising these trades API-server load and controller CPU for lower queue latency. Because 3.2 and 3.3 changed hydrator behavior, if you use it, size hydration processors after upgrading rather than before.

Signals to compare before and after

Capture a baseline on the old version, then compare the same signals on the new one:

Signal Where What a regression looks like
Reconcile queue depth controller work-queue metrics Depth rising after restart and not draining
Application count by sync status argocd_app_info labels Sudden OutOfSync spike right after upgrade
Git request latency repo-server metrics Longer manifest generation after Helm or Kustomize bump
Kubernetes API request rate apiserver audit or metrics Higher list/watch volume if exclusions were overridden
Redis memory Redis exporter Growth if health and cache behavior changed

Metric names are listed generically here because they shift between releases (3.0 removed three, and 3.4’s OpenTelemetry bump can rename others). Verify names against your running /metrics endpoint.

A Rollback-Safe Rollout Plan

A defensible Argo CD upgrade guide for production has a checkpoint before every irreversible step. Argo CD makes this easier than most control planes because the desired state lives in Git, not in Argo CD. What you must protect is the Argo CD configuration itself, the CRD versions, and the credentials.

Phase 0: inventory and backup

  1. Record the running version and how it was installed (raw manifests, Kustomize, Helm, or Argo CD managing itself).
  2. Export settings using the disaster recovery tool. The documented form runs argocd admin export from the matching image tag and argocd admin import - to restore. If Argo CD lives outside the argocd namespace, pass -n; the docs warn the export will not error if pointed at the wrong namespace, producing an incomplete backup silently. Check the file for your Applications, AppProjects and Secrets before proceeding.
  3. Snapshot the ConfigMaps (argocd-cm, argocd-rbac-cm, argocd-cmd-params-cm) and Secrets into an encrypted store. Store them separately from the export.
  4. Run the repository-in-argocd-cm detection command from the upgrade guide and fix findings.

Phase 1: convert silent changes into explicit settings

Before changing the version, commit explicit values for anything you want to keep stable. This is safe on 2.14 because the keys are accepted, and it makes the upgrade diff readable:

# argocd-cm  (pin v2 behaviour during the upgrade window)
data:
  application.resourceTrackingMethod: label
  server.rbac.disableApplicationFineGrainedRBACInheritance: "false"
  resource.compareoptions: |
    ignoreResourceStatusField: crd
    ignoreDifferencesOnResourceUpdates: false
---
# argocd-cmd-params-cm
data:
  controller.resource.health.persist: "true"

Then pre-write the logs, get policies and the rewritten Dex subjects. I recommend verifying key names against the version you install, as I derived these from the 3.0 upgrade page and did not test them on a live 2.14 instance. If a key is unknown on 2.14, keep it in a change file that you apply with the upgrade.

Phase 2: staging, then a canary tenant

Upgrade a staging instance, ideally pointed at real repositories in read-only or no-auto-sync mode. Run the manifest diff loop, run the RBAC assertions, and confirm SSO login with a real Dex user. Only then plan production.

For production, if you run one Argo CD per environment or per region, upgrade the least critical instance first. If you run a single shared instance, disable automated sync and prune on a canary set of applications for the window, then re-enable them after verifying.

Phase 3: apply, verify, decide

Apply the full manifest with server-side apply and force-conflicts:

kubectl apply -n argocd --server-side --force-conflicts \
  -f https://raw.githubusercontent.com/argoproj/argo-cd/v3.5.2/manifests/install.yaml

Use the HA manifest path if you run HA, and vendor the manifest into Git instead of applying from a URL where your change process requires it. After the controller comes up, verify in this order: pods ready, login works, Application count matches the pre-upgrade count, no unexpected OutOfSync jump, logs visible for a test role, alerts still firing (fire a synthetic one), then sync one application under each tracking method.

Phase 4: rollback triggers

Define them before you start. Reasonable triggers: more than a small, pre-agreed fraction of applications flipping to OutOfSync that you cannot explain; SSO login failure; repo-server crash loops; metrics or alerts silent. Rollback means re-applying the previous manifest and restoring the config snapshot. There is an important caveat: the official upgrade overview does not describe a formal downgrade procedure, so I treat rollback as “restore the prior install and settings from backup,” not as a supported reverse migration. The safest hedge is that Argo CD’s state of record is Git, so a restore mostly re-derives from Git, but any Application CR edits made during the window would be lost.

Then, only after a stable observation period, remove the pinned v2 settings one at a time, one change per window.

Trade-offs, Gotchas, and What Goes Wrong

Pruning after selector changes. If your ApplicationSets now generate fewer Applications because nested selectors are always applied, and an ApplicationSet policy allows deletion, Argo CD may delete the Applications that fell out of the set, cascading to their workloads if finalizers are set. Before upgrading, set ApplicationSet sync policy to preventing deletes, or verify generated sets in staging. This is my strongest single warning.

Sync waves and tracking switches. Changing tracking method with automated pruning on can cause resources that lack the new annotation to be considered untracked in edge cases, especially where a resource has two owners. Keep pruning off for the first sync after the change, review the diff, then re-enable it.

Resource exclusions hide things you manage. Exclusions are global. If a team deploys a Lease or an Endpoints object through Argo CD on purpose, it will vanish from the tree. Override resource.exclusions explicitly and document it.

Dex identities. Groups still work, but user-specific RBAC entries can fail silently, leaving people with the default role. Test with real accounts, not the admin user.

Deprecated but present. Several 3.x items are deprecated rather than removed: the v1 actions endpoint, --self-heal-backoff-cooldown-seconds, .spec.signatureKeys and the verifyResult field. Treat deprecation warnings in logs as a backlog, not noise. Grep controller and server logs for “deprecat” weekly.

Skipping the CRD step. From 3.3, applying the ApplicationSet CRD client-side fails with an annotation-size error; the fix is server-side apply. If a GitOps pipeline applies Argo CD’s own manifests, update it before the version bump or your bootstrap will break at the worst moment.

Change bundling. Bundling an Argo CD upgrade with a Kubernetes upgrade, a Helm chart bump and a Dex change is how a two-hour window becomes two days. Argo CD tracks Kubernetes versions via the cluster version string (changed in 3.4), so separate the events.

Where this advice is weakest. I have relied on the official upgrade pages; the project cannot document every plugin, extension or in-house automation. UI extensions built on React 18 need a rebuild for 3.5, and gRPC clients need regenerating. If you own any of those, they are your real critical path, and no guide can size that effort for you.

Practical Recommendations

If you take one idea from this guide, take this: upgrade in two separate motions, first the binaries with v2 behavior pinned, then the behavior changes one at a time. That sequencing costs an extra maintenance window and buys you attribution, which is what you need when something breaks at 2 a.m.

Choose your target deliberately. For a new upgrade today I would land on the newest patch of 3.4 or 3.5 rather than 3.0, because 3.0 through 3.2 are out of the patch window and because you will read the intermediate guides regardless. Prefer 3.4 if you have UI extensions, gRPC clients or plain-HTTP OCI registries that the 3.5 changes would touch; prefer 3.5 if you want the newest supported line and the Source Integrity and mTLS features. This is judgment, not project guidance.

Checklist for the change ticket:

  • [ ] Backup exported with the matching image tag and verified non-empty for Applications, AppProjects and Secrets
  • [ ] Repository definitions moved out of argocd-cm into Secrets
  • [ ] Explicit logs, get and update/*, delete/* policies written and validated with argocd admin settings rbac can
  • [ ] Dex user subjects rewritten, tested with real accounts
  • [ ] Dashboards and alerts migrated from removed metrics to argocd_app_info
  • [ ] ApplicationSets diffed in staging; deletion protection on during the window
  • [ ] Tracking method decided (annotation+label as the stepping stone), installationID set for multi-instance clusters
  • [ ] Server-side apply used for manifests; self-managing Application has ServerSideApply=true
  • [ ] Sharding algorithm and replica count unchanged during the version window
  • [ ] Rollback triggers agreed and the previous manifest and config snapshot at hand
  • [ ] Plan to remove pinned v2 settings, one per window, with owners

Frequently Asked Questions

What are the main Argo CD v3 breaking changes?

The 3.0 release changed fine-grained RBAC so update and delete no longer inherit to sub-resources, made log access always RBAC-checked, switched resource tracking to annotations, added default resource exclusions, removed three legacy metrics, moved Dex identity to federated_claims.user_id, removed repositories from argocd-cm, ignored applyNestedSelectors, and bundled Helm 3.17.1. Each has a documented config key to restore v2 behavior temporarily.

Can I upgrade straight from Argo CD 2.14 to 3.5?

The project’s rule is that minor releases may need workarounds and you should read each intermediate upgrade guide. Technically you can apply a newer manifest directly, but you inherit every change in between, including server-side apply for the ApplicationSet CRD from 3.3. Test in staging first, and prefer landing on a supported patch of 3.3, 3.4 or 3.5 rather than stopping at 3.0.

Will Argo CD 3 upgrade restart or delete my workloads?

Upgrading the control plane does not restart workloads, because Argo CD is not in the data path. The real risks are indirect: automated sync with prune enabled acting on changed behavior, such as fewer generated Applications from an ApplicationSet, or tracking changes causing resources to be reconsidered. Disable automated prune during the window and enable deletion protection on ApplicationSets.

Do I have to switch to annotation tracking?

No. Set application.resourceTrackingMethod to label to keep v2 behavior, or annotation+label to write annotations while keeping the instance label for other tools. Annotation tracking is the default because labels cap at 63 characters and collide with other tools. Plan to migrate eventually, and set installationID if several Argo CD instances share a cluster.

How does Argo CD sharding change in version 3?

The sharding model is the same: controller replicas own clusters, set via ARGOCD_CONTROLLER_REPLICAS and controller.sharding.algorithm with legacy, round-robin and consistent-hashing options, the latter two still documented as experimental. The 3.0 defaults reduce event volume, which may lower the replicas you need. Do not change algorithm and version in one window.

How do I roll back a failed Argo CD upgrade?

The official overview does not document a formal downgrade. In practice, re-apply the previous version’s full manifest and restore the exported settings and ConfigMap snapshots, noting that any Application changes made during the window may be lost. Because desired state lives in Git, most state re-derives. Define rollback triggers ahead of time and test the restore path in staging.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *