Istio 1.31 Explained: What Changed and How to Upgrade Your Service Mesh
Most service mesh upgrades fail for boring reasons: a default flipped quietly, an artifact registry moved, or a restart wave hit production before anyone watched the canary. Istio 1.31 has all three ingredients. It ships weighted waypoint canaries for ambient mode, mesh-wide default traffic policies, zone-aware load balancing, a FIPS 140-3 compliance policy, and a behavior change in how unhealthy endpoints are sent to proxies. It also ends publication of release artifacts to Google Cloud registries, which will break pipelines that still pull Helm charts or binaries from the old locations.
This matters now because Istio 1.29 reaches end of life on 12 October 2026, according to the project’s supported-releases page, so many teams have a forcing function this month. The 1.31 release is also the first where ambient mode gets a first-class progressive-delivery story for its Layer 7 waypoints.
This post walks through what actually changed, which changes can hurt you, and a revision-based canary upgrade with rollback that works for sidecar, ambient, and mixed meshes.
What this covers: the verified 1.31 change list, the ambient architecture it builds on, upgrade-note landmines, exact canary and rollback commands, a sidecar-versus-ambient decision framework, and an FAQ.
Context and Background
Istio has been the reference implementation of the service mesh pattern since 2017: a control plane (istiod) distributes configuration and workload identity, and a data plane of proxies enforces mutual TLS, routing, authorization, and telemetry. For most of its life the data plane meant an Envoy sidecar injected into every pod. That model is powerful but expensive. Every pod carries a proxy, every proxy needs CPU and memory, and every Istio upgrade means restarting every workload to pick up the new proxy version.
Ambient mode, introduced as an alternative data plane and declared stable in the 1.24 release line, splits the proxy into two layers. A per-node Rust proxy called ztunnel handles Layer 4 concerns: mutual TLS, identity, L4 authorization, and basic telemetry. Optional Envoy-based waypoint proxies, running as ordinary deployments outside application pods, add Layer 7 features only where a team asks for them. The official overview states that ztunnel does not parse HTTP at all, and that workloads in different data plane modes interoperate inside the same mesh. That interoperability is what makes incremental migration possible.
Where do the competitors sit? Linkerd keeps a deliberately small sidecar proxy written in Rust, and Cilium runs mesh functions in the kernel and in per-node Envoy. We compared the trade-offs in depth in our ADR on Istio ambient mesh versus Linkerd and in the Cilium service mesh versus Istio ambient ADR. If you are still choosing a mesh, read those first. If you already run Istio, the question for this release is narrower: what do I have to do, what can break, and what do I get?
The Istio project’s own record is the primary source for everything below. The facts in this post come from the Istio 1.31.0 announcement, its upgrade notes and change notes, and the supported-releases and canary-upgrade documentation, all read in October 2026. Where the project’s pages disagree with each other, or where something is not documented, this post says so.
A word on release timing. The announcement page lists 31 August 2026 as the 1.31.0 release date, while the supported-releases table shows 27 August 2026. We could not reconcile the two, so we treat “late August 2026” as the verified claim. The documentation site was already serving 1.31.1 pages when we checked, which means at least one patch release exists. Always install the latest patch of the minor version rather than 1.31.0 itself.
Support windows matter for planning. Per the supported-releases page, 1.31 is the latest release with an end of life of roughly February 2027, 1.30 runs to roughly December 2026, and 1.29 ends on 12 October 2026. The project’s policy is to support a minor release until six weeks after the second subsequent minor release ships. In practice, that gives you about two quarters per minor, so skipping versions is a bet on your own calendar.
What Istio 1.31 Actually Changes
The short version for a skimming reader: 1.31 is an incremental release with a few operationally significant defaults. It supports Kubernetes 1.32 through 1.36, uses Envoy release 1.39 as its data plane base according to the supported-releases page, and the version listing shows no known CVEs against 1.31.0 and later at the time of checking. The headline additions fall into four groups, shown in the figure below.

Figure 1: The Istio 1.31 change surface, grouped by area. Source: Istio 1.31.0 announcement, change notes, and upgrade notes.
The diagram is a map of the release, not a priority order. Read the long-description version: ambient gains weighted waypoint canaries and an agentgateway waypoint class; traffic management gains zone-aware load balancing and a mesh-wide default policy; security gains trust-domain matching in authorization and a FIPS profile; and the one hard deprecation is the end of artifact publication to Google Cloud locations.
Ambient mode: waypoint canaries and agentgateway
The most interesting ambient change is weighted waypoint canaries. A waypoint is the Layer 7 proxy that sits in front of a namespace or a service in ambient mode. Until now, upgrading or reconfiguring a waypoint was effectively an all-or-nothing event for the workloads it served. The 1.31 release adds configurable traffic splitting between a primary waypoint and a canary waypoint, so you can send, for example, a small fraction of traffic through a waypoint running a new Envoy version or a changed policy set before cutting over.
This is the first release where the waypoint, which is the only ambient component that resembles a traditional proxy fleet, can be rolled out progressively with the mesh itself doing the weighting. The change notes list it as a feature but do not publish the exact configuration surface in the summary we reviewed. Check the 1.31 docs for the precise resource fields before depending on it. We do not reproduce field names we did not verify.
Second, the release adds an istio-agentgateway-waypoint GatewayClass, which lets an agentgateway data plane act as a waypoint. Agentgateway first appeared as an experimental gateway class in Istio 1.30, described as a proxy built for AI agent and Model Context Protocol (MCP) server traffic. The 1.31 announcement extends that to the waypoint role. Both the 1.30 and 1.31 materials position this as new, and the project does not claim production maturity in the pages we read. Treat it as an opt-in experiment for now.
Third, 1.31 includes multicluster stability fixes in ambient mode, covering credential rotation, memory leaks, and issues in the CNI node agent. If you run ambient across clusters, those fixes are reason enough to move. We examined the multicluster pattern and its gateway implications in our piece on Istio ambient multicluster with the Gateway API Inference Extension for LLM serving.
Two smaller ambient items are worth noting. The change notes mention optimized xDS pushes in ambient mode, controlled by an AMBIENT_SCOPED_ADDRESS_PUSHES environment variable, which scopes address updates so istiod does less work in large meshes. Separately, the 1.30 release introduced a migration guide from sidecar to ambient, which remains the best starting point for namespace-by-namespace moves.
Traffic management: zone-awareness and default policies
Zone-aware load balancing lets the mesh prefer endpoints in the caller’s availability zone and fail over to other zones when local capacity or health drops. The setting appears as zoneAwareLbSetting in the change notes. The practical win is cost and latency: cross-zone traffic is billed on most public clouds and adds round-trip time. The practical risk is imbalance. If one zone hosts far more callers than backends, local preference can overload the local pods while remote pods sit idle. Zone preference needs per-zone replica counts that match caller distribution.
defaultTrafficPolicy in MeshConfig sets mesh-wide baseline behavior such as connection pooling and outlier detection. Before this, getting a sane default meant creating a DestinationRule per host or per namespace. The new field makes it possible to define a baseline once and override it where needed. The change notes do not say how precedence resolves against an existing DestinationRule in the summary we reviewed, so test the interaction in staging before assuming a rule overrides the default the way you expect.
Other additions: a prefix_rewrite field for HTTP redirect path rewriting, a budget_interval setting for retry budgets, meshConfig.serviceEntryVisibility for controlling ServiceEntry reach, a new ALLOW_ANY_DYNAMIC_DNS outbound policy mode that supports a dynamic forward proxy for unknown hosts, sidecar egress host exclusion using a ~ prefix, and configurable HTTP/2 keepalive PING settings on upstream connections. The dynamic forward proxy mode deserves care. It resolves arbitrary DNS names at request time, which is convenient for egress to SaaS endpoints and also removes the allow-list property that REGISTRY_ONLY gives you.
Security: trust domains, FIPS, and tighter control-plane access
The trustDomains and notTrustDomains fields in AuthorizationPolicy let you match requests by the SPIFFE trust domain of the caller. That sounds minor until you run multiple clusters or multiple meshes with different trust domains and want a policy such as “only accept calls from trust domain A for this workload.” It closes a gap where you previously had to enumerate principals or namespaces.
The FIPS 140-3 compliance policy is selected with COMPLIANCE_POLICY=fips-140-3. Per the change notes it enforces TLS 1.2 or later with compliant cipher suites and auto-injects GODEBUG=fips140=only for Go components. Important caveat: a policy that restricts cipher suites is not by itself a FIPS validation of your whole deployment. Whether your build, your Envoy binary, and your crypto module satisfy an auditor is a separate question the release notes do not answer.
Two hardening changes are easy to miss. First, strict gateway merging (PILOT_ENABLE_STRICT_GATEWAY_MERGING) prevents merging of Gateway resources across namespaces. Second, the MCP config-serving endpoint on istiod now requires a verified control-plane identity. This ties into the upgrade notes below, because it changes behavior for custom consumers.
Istio 1.30 had already made debug endpoint authentication on port 15010 mandatory by default and made the control-plane TLS minimum version configurable. If you skipped 1.30, you inherit that behavior change along with 1.31’s.
The artifact deprecation
The one explicit deprecation in 1.31 is not an API. From 1.31 onward, the project no longer publishes artifacts to gcr.io/istio-release and related Google Cloud locations. Docker images stay on Docker Hub. Helm charts and other artifacts move to blob.istio.io/istio-release and ghcr.io/istio/release/charts. The project is running four “scream tests” that temporarily disable the GCP-hosted artifacts between September and December 2026, the last spanning 8 and 9 December 2026.
A scream test deliberately breaks something to find who still depends on it. If your CI, air-gapped mirror, or GitOps repository references the old locations, you will find out on someone else’s schedule. We cover the audit for this in the upgrade section below.
How the Ambient Data Plane Works, and Why 1.31 Upgrades Differ
An Istio upgrade is not one operation. It is at least four, and which ones apply depends on your data plane. The control plane (istiod) upgrades like any Kubernetes deployment. Sidecars upgrade only when pods restart. In ambient mode, ztunnel runs as a DaemonSet and the CNI node agent runs alongside it, so both upgrade per node. Waypoints are Deployments and upgrade like any other workload, with the new weighted canary available as a safety net.

Figure 2: A mixed Istio mesh. Ambient pods ride ztunnel and optionally a waypoint over HBONE, while sidecar pods keep their Envoy. Source: Istio ambient documentation, redrawn.
The figure shows why mixed meshes are the norm during migration. The same istiod serves xDS configuration to sidecars and waypoints, and issues workload certificates to ztunnels. Traffic between an ambient pod and a sidecar pod is tunneled over HBONE (HTTP-based overlay network environment), an HTTP CONNECT tunnel carrying mutual TLS.
The direct answer: what an upgrade has to touch
For an Istio 1.31 upgrade, you must upgrade istiod, then each data plane component you run: sidecar proxies via pod restart, ztunnel and the CNI node agent via DaemonSet rollout in ambient mode, and waypoint proxies via deployment rollout. The control plane may run one minor version ahead of the data plane, never behind. That one-version skew rule is what makes a staged, revision-based upgrade safe.
Control plane first, always
The supported-releases page states that the control plane can run one version ahead of the data plane but not the reverse. This is the load-bearing constraint of every Istio upgrade. If you upgrade the data plane first, sidecars may speak xDS features the old istiod does not understand. If you upgrade the control plane first and restart nothing, existing sidecars keep working against the newer istiod within the skew window.
This asymmetry has a planning consequence. A control plane upgrade by itself is low risk and reversible, while the restart wave that follows is where incidents happen. A good upgrade plan therefore separates the two events by hours or days, not minutes. You install the new istiod, confirm it is healthy and idle, then move workloads in waves.
Why ambient changes the blast radius
In a sidecar mesh, a proxy upgrade is a per-pod event. A bad Envoy version breaks the pods that restarted, and the blast radius grows with each wave of restarts. You control it by wave size.
In ambient mode the unit is different. A ztunnel runs per node and carries the L4 traffic for every ambient pod on that node. Upgrading ztunnel on a node briefly affects all those pods at once, which is a larger blast radius per node but a much smaller number of proxies overall. In exchange, the application pods themselves never restart, so there is no rolling restart of the workload fleet and no coordination with application owners.
Waypoints sit between these two models. They carry L7 policy for a whole namespace or service, so an upgrade affects every consumer of that waypoint. The 1.31 weighted canary exists to shrink that blast radius, letting you shift a slice of traffic first.
The Istio ambient documentation indexes a Helm-based upgrade guide for ambient installs, which is the supported path for ztunnel, CNI, and istiod together. The landing page we reviewed did not spell out the procedure, so follow that guide for exact commands. This post gives the revision-based flow for the control plane and sidecars in full, because that is documented precisely.
What the sidecar-to-ambient migration guide adds
Istio 1.30 introduced a dedicated sidecar-to-ambient migration guide. The sensible reading for 1.31 planning is that you do not need to pick one mode for the whole fleet. A namespace is enrolled in ambient by label. You can move one namespace, watch it, and move it back. During that period your mesh is mixed, and both data planes talk to each other.
Two failure patterns recur in mixed meshes. The first is policy that was written for sidecars and silently behaves differently on ambient, because L7 policy is only enforced where a waypoint exists. A namespace moved to ztunnel-only will keep enforcing L4 authorization but stop enforcing HTTP-level rules until a waypoint is attached. The second is observability drift: dashboards built on sidecar metrics will not show the same series for ambient workloads, because L7 metrics come from waypoints and L4 metrics from ztunnel.
Upgrade Notes That Can Hurt: Read These Before You Restart Anything
The 1.31 upgrade notes list six items. Each one is a behavior change you can verify before it reaches production. We rank them here by how likely they are to surprise a typical platform team, and we state the mitigation the project documents.
Unhealthy endpoints are now sent by default
The most important behavior change is this one. Per the upgrade notes, Istio now sends unhealthy endpoints to proxies by default unless OutlierDetection.minHealthPercent is configured. It can be turned off with the PILOT_AUTO_SEND_UNHEALTHY_ENDPOINTS environment variable or via compatibility profiles.
Why does that matter? Historically, istiod filtered unhealthy endpoints out of the configuration it pushed to proxies. Sending them lets the proxy make health-aware decisions, such as panic-threshold behavior, using full information. The consequence is that your proxies receive a larger endpoint set during incidents, and load balancing behavior during partial outages can differ from what your runbooks describe. If you rely on the old behavior, set the flag explicitly and plan to revisit it.
The upgrade notes do not quantify the performance impact, and neither do we. Compare proxy memory and xDS push sizes in staging under a simulated pod failure before you rely on either setting.
WorkloadEntry over HBONE needs a label
Previously auto-registered WorkloadEntry resources need either re-registration or a manual networking.istio.io/tunnel=http label to keep working over HBONE tunneling after the upgrade. This affects teams that bridge virtual machines or non-Kubernetes workloads into an ambient mesh. If you have none, skip it. If you do, list your WorkloadEntries before the upgrade and relabel them in the same change window, otherwise those workloads can lose connectivity through the tunnel.
Larger reconnect messages in very large meshes
The WDS (workload discovery service) reconnect request now includes workload versions. In large meshes of roughly 40,000 or more workloads, per the notes, the request can exceed istiod’s default 4 MiB gRPC receive limit. The documented mitigation is raising ISTIO_GPRC_MAXRECVMSGSIZE. The spelling in the notes is exactly as written there, so copy it from the source rather than correcting it. Most clusters are far below this scale, but multicluster ambient deployments that aggregate workloads can approach it.
Authentication on the xDS API generator
Custom MCP consumers now need a control-plane identity when connecting to istiod’s api generator from non-system namespaces. Setting ENABLE_XDS_API_GENERATOR_AUTH=false restores the old behavior. Most teams never touch this, but if you run a custom config consumer or a third-party tool that reads from istiod’s API generator, test it. Disabling the new authentication is a security regression, so treat it as a temporary bridge.
A removed feature flag
The PILOT_SPAWN_UPSTREAM_SPAN_FOR_GATEWAY environment variable is gone, and the behavior it controlled is now always on. This affects gateway tracing span structure. If you set the variable in manifests, remove it, because the flag no longer exists and tracing dashboards may show an additional upstream span.
The registry move
The artifact change described earlier is the one most likely to hit CI first. Audit anything that references the old locations: Helm repository URLs, container image mirrors, istioctl download scripts, and Terraform modules. Switch Helm sources to ghcr.io/istio/release/charts or blob.istio.io/istio-release. Because the scream tests run through December 2026, an unmigrated pipeline might pass today and fail on a test day, which is the worst kind of intermittent failure.
A practical pre-flight
Before touching production, run this sequence on a staging cluster that mirrors your real configuration. First, run istioctl analyze with the 1.31 binary, since 1.31 adds checks for ServiceEntry protocol conflicts and Gateway API CRD version validation. Second, capture baseline p99 latency, error rate, istiod CPU and memory, and xDS push counts. Third, install the new revision (next section) and compare. Fourth, simulate a pod failure and confirm that load balancing under the new unhealthy-endpoint behavior matches your expectations.
The Revision-Based Canary Upgrade, Step by Step
Revision-based upgrades install the new control plane next to the old one instead of replacing it. Each istiod carries a revision name, and each namespace selects exactly one revision through a label. Nothing moves until you relabel a namespace and restart its pods, which makes every step reversible. The Istio canary upgrade documentation recommends revision tags as stable aliases, and that is the pattern below.

Figure 3: The revision-based canary upgrade for Istio 1.31. Rollback at any step before retirement is a tag change and a restart. Source: Istio canary upgrade documentation, adapted.
The sequence runs left to right in time. You install the new revision, shift one canary namespace, observe, then move the production tag in waves, and only then retire the old control plane.
Step 1: Take inventory and install the new revision
Start by recording what you have. List installed revisions and which namespaces use which, so you know what “rollback” means for your cluster.
# What control planes and tags exist?
istioctl tag list
kubectl get mutatingwebhookconfigurations | grep istio
# Which namespaces use which revision?
kubectl get ns -L istio.io/rev -L istio-injection -L istio.io/dataplane-mode
Download the 1.31 istioctl (from a location that is not the deprecated registry) and install the new control plane under a revision name. The documentation example uses canary; we recommend a version-encoded name such as 1-31-1 so the revision is self-describing.
istioctl version # confirm you are using the 1.31 binary
istioctl x precheck # check cluster readiness
istioctl install --set revision=1-31-1 -y
kubectl get pods -n istio-system -l app=istiod
Two istiod deployments now run side by side. Existing workloads still use the old one, so nothing has changed for them. This is the cheapest moment to catch a bad install, so verify the new istiod is ready, its logs are clean, and its webhook exists.
Step 2: Create revision tags
Tags decouple namespace labels from version numbers. Without tags you must relabel every namespace on every upgrade. With tags you relabel once and then repoint the tag.
# Existing production points at the old revision
istioctl tag set prod-stable --revision 1-30-1
# The candidate points at the new revision
istioctl tag set prod-canary --revision 1-31-1
# One low-risk namespace opts into the canary tag
kubectl label namespace shop-canary istio.io/rev=prod-canary --overwrite
kubectl label namespace shop-canary istio-injection- # remove old label if present
# Production namespaces use the stable tag
kubectl label namespace shop istio.io/rev=prod-stable --overwrite
The documentation uses exactly this pattern with prod-stable and prod-canary names and with 1-30-1 and 1-31-1 as revisions. Keep the istio-injection label off namespaces that use istio.io/rev, because the two conflict and the older label can win.
Step 3: Restart the canary and observe
Sidecars pick up the new proxy version only when pods are recreated.
kubectl rollout restart deployment -n shop-canary
istioctl proxy-status | grep shop-canary # confirm they connect to 1-31-1
Verify the data plane version and sync state. istioctl proxy-status shows each proxy with the istiod it is connected to and whether configuration is acknowledged. A canary pod stuck in a non-synced state means the new control plane and the proxy disagree about configuration, which is exactly what you wanted to learn before production.
What should you watch during the canary window? Compare against the baseline you captured earlier: request error rate by response code, p50 and p99 latency, TLS handshake failures, istiod CPU, memory, and xDS push latency, and the number of endpoints per cluster in a representative proxy. The last one is the early indicator for the unhealthy-endpoint change, because that default alters what endpoints a proxy receives. Hold the canary long enough to include a deployment event and, if feasible, a controlled pod kill, since steady-state traffic hides failure-path behavior.
Step 4: Promote in waves
When the canary is clean, repoint the stable tag and restart in waves. This is the documented promotion step.
istioctl tag set prod-stable --revision 1-31-1 --overwrite
kubectl rollout restart deployment -n shop
The advantage of repointing a tag is that namespaces already labeled prod-stable pick up the new revision at their next restart, with no relabeling. The danger is the same fact: any pod that restarts for any reason, for example a node drain at 3 a.m., now gets the new proxy. Control the blast radius by deciding when to repoint relative to your normal deploy cadence, and consider keeping a second tag for namespaces you want to migrate later.
For a large fleet, do it in waves by tier: internal tools first, then low-traffic services, then the critical path. Pause between waves long enough to see a full cycle of your usual traffic pattern. A rolling restart across hundreds of deployments at once defeats the purpose of the canary.
Step 5: Rollback
Rollback is the reason to use revisions. Before you retire the old control plane, rolling back is symmetrical to rolling forward.
istioctl tag set prod-stable --revision 1-30-1 --overwrite
kubectl rollout restart deployment -n shop
If the problem is in the canary only, the documentation’s rollback is to uninstall the canary revision, and you then have to reinstall old gateways manually if you used in-cluster gateways on the new revision.
istioctl uninstall --revision 1-31-1 -y
Be honest about the limit here: once the restart wave has recreated pods on the new proxy, rolling back means restarting them again. Revision-based rollback is fast for the control plane and slow for the data plane, because the data plane moves at the speed of your restarts.
Step 6: Retire the old revision
Only after the new revision has carried production traffic through a full business cycle, remove the old control plane.
istioctl uninstall --revision 1-30-1 -y
Check that no namespace still points at the old revision and that no tag references it before you do this. Pods still connected to the old istiod lose their control plane when it is deleted. The “one version ahead” rule means you should also schedule the next upgrade promptly, because a mesh frozen at 1.30 data plane with a 1.32 control plane would violate the skew limit.
Upgrading ambient components
For ambient installs, the equivalent unit is the Helm release set: base, istiod, CNI, and ztunnel, and optionally gateway charts for waypoints. Use the ambient Helm upgrade guide, upgrade istiod first, then the CNI and ztunnel DaemonSets with a conservative maxUnavailable so only a fraction of nodes are affected at once. Because ztunnel upgrades interrupt L4 connections on that node, drain or stagger nodes if your workloads cannot tolerate connection resets.
For waypoints, use the 1.31 weighted canary: deploy a canary waypoint on the new proxy version, send a small weight to it, watch L7 error rates, then raise the weight. The exact resource fields are in the 1.31 documentation; we have not verified them here, so we do not quote them.
Also note the Helm detail from 1.30: its upgrade notes mention webhook failurePolicy changes during Helm v4 upgrades. If you are on Helm v4, read that note, because a webhook with the wrong failure policy can block pod creation during the upgrade.
Sidecar or Ambient: A Decision Framework for This Release
Istio 1.31 does not force a mode decision, but it makes the ambient side more attractive in two ways: waypoint canaries reduce the L7 upgrade risk, and multicluster fixes reduce the operational risk. It also leaves the sidecar side fully supported. The right answer depends on your workloads, not on the release.

Figure 4: A decision flow for picking a data plane per namespace, not per cluster. Source: author analysis based on Istio ambient documentation.
The flow is an opinion, not an official rubric. Its logic: if you rely heavily on per-pod Envoy customization, such as EnvoyFilter and Wasm, sidecars give you the most control. If you need strict per-pod isolation of the proxy and its credentials, a sidecar gives each pod its own proxy identity path, while ztunnel is a shared per-node component. If neither applies and sidecar resource cost is a real line item, ambient is the natural fit. Everything else lands in a mixed mesh where you migrate namespace by namespace.
A decision matrix
| Factor | Sidecar | Ambient (ztunnel only) | Ambient plus waypoint |
|---|---|---|---|
| L4 mTLS and identity | Yes | Yes | Yes |
| HTTP routing and L7 authorization | Yes, in pod | No | Yes, in waypoint |
| Upgrade requires pod restart | Yes | No for app pods | No for app pods |
| Upgrade blast radius | Per pod | Per node | Per waypoint consumer set |
| Per-pod proxy resource cost | One Envoy per pod | None | None, shared waypoint |
| EnvoyFilter and Wasm customization | Full | Not applicable | Limited to waypoint |
| Maturity of tooling and docs | Longest history | Stable since 1.24 | Newer, canary support in 1.31 |
Read the matrix as a starting point for conversation. The “limited to waypoint” cell is a judgment from how the architecture works, since L7 extensions can only run in the proxy that handles L7. Check your actual extension inventory before migrating.
One more factor deserves mention: the data plane mode determines the upgrade operation itself. A sidecar mesh upgrade is dominated by restart coordination with application owners. An ambient upgrade is dominated by node-level DaemonSet rollouts that the platform team controls alone. If your organization finds cross-team restart coordination to be the bottleneck of every Istio upgrade, that organizational cost alone is a legitimate reason to move namespaces to ambient.
Trade-offs, Gotchas, and What Goes Wrong
Upgrading the control plane is easy; restarting the fleet is not. Teams often declare victory when the new istiod is healthy. The real risk begins at the first restart wave. Plan for pods that restart unexpectedly during the canary window, because they will join the revision their namespace tag points to. If a tag is repointed mid-incident, a routine node drain becomes an unplanned proxy upgrade.
Skipped minors compound risk. Each minor release carries behavior changes, and the project’s support policy offers roughly two quarters of life per minor. Going from 1.29 straight to 1.31 means you absorb both the 1.30 changes (debug endpoint authentication on port 15010 by default, configurable control-plane TLS minimum) and the 1.31 changes at once. Revision-based upgrades tolerate this, but your debugging does not, since two sets of changes obscure the cause of any regression. Upgrade one minor at a time if you can.
Defaults changed, not just features. The unhealthy-endpoint behavior is a default flip, and default flips are the dangerous kind because no manifest of yours changes. Diff your effective mesh config before and after. Check the compatibility profile mechanism in the documentation if you need to pin old behavior while you test.
Zone-aware load balancing can concentrate load. Local preference is only as good as the match between caller distribution and backend capacity per zone. A deployment with three replicas across three zones and uneven caller traffic can saturate a single zone. Pair the feature with outlier detection, per-zone autoscaling, and a failover test.
The dynamic forward proxy trades control for convenience. ALLOW_ANY_DYNAMIC_DNS helps when applications call many unknown external hosts, but it weakens the egress allow-list you get from REGISTRY_ONLY. If a compliance regime requires you to enumerate permitted destinations, do not use it.
FIPS policy is not FIPS certification. The fips-140-3 compliance policy sets TLS and cipher constraints and an environment for Go components. It does not certify your entire stack. If an audit depends on it, get the validated-module claims for your specific Envoy build and OS from your vendor, because the release notes are silent on that.
Ambient is not free of operational cost. You trade per-pod proxies for node-level components. A misbehaving ztunnel affects every ambient pod on its node. The CNI node agent needs privileged access to program traffic redirection and must coexist with your cluster’s primary CNI. Test with the CNI you actually run, including any chaining configuration.
Agentgateway classes are new. Both 1.30 and 1.31 materials introduce agentgateway as new, and the pages we read do not label it production-ready. We could not find published performance comparisons against Envoy in the Istio materials, so any claim about its throughput or latency in your environment must come from your own testing.
Do not trust a single signal. A canary that looks healthy on error rate may hide a slow regression in tail latency or control-plane load. Include istiod resource use and xDS push times in the gate, not only application metrics.
Practical Recommendations
Treat 1.31 as a planned, two-stage change: first get the artifacts and defaults right, then move the data plane. Most of the risk sits in preparation rather than in the install commands, which are short.
If you are on 1.29, move now. Its end of life is 12 October 2026, and after that date a critical CVE will not get a patch for your minor. If you are on 1.30, you have until roughly December 2026, which still argues for an upgrade this quarter. Aim for the latest 1.31 patch, since the docs already show 1.31.1.
Move one namespace to ambient only after you have inventoried L7 policy and extensions. For anything that depends on EnvoyFilter, Wasm, or HTTP-level authorization, either keep sidecars or attach a waypoint first and confirm the policy still applies.
Use this checklist as your change ticket:
- Audit CI, Helm sources, mirrors, and download scripts for
gcr.io/istio-release; repoint to Docker Hub images,ghcr.io/istio/release/charts, orblob.istio.io/istio-release. - Run
istioctl analyzeusing the 1.31 binary on every cluster, and fix new ServiceEntry or Gateway API CRD findings. - Decide on
PILOT_AUTO_SEND_UNHEALTHY_ENDPOINTSexplicitly, and test load balancing under pod failure. - Relabel any auto-registered
WorkloadEntrywithnetworking.istio.io/tunnel=httpif you use HBONE for VMs. - Install the new revision beside the old; create
prod-stableandprod-canarytags. - Canary one low-risk namespace; hold through a deploy and a pod-kill test.
- Promote in tiered waves, pausing for a full traffic cycle between them.
- Keep the old revision for a full business cycle, then uninstall it.
- Record rollback commands in the change ticket before you start.
Frequently Asked Questions
What is new in Istio 1.31?
Istio 1.31 adds weighted waypoint canaries and an agentgateway waypoint GatewayClass for ambient mode, zone-aware load balancing, a mesh-wide defaultTrafficPolicy, trustDomains matching in AuthorizationPolicy, a FIPS 140-3 compliance policy, and a dynamic forward proxy outbound mode. It supports Kubernetes 1.32 to 1.36 and stops publishing artifacts to Google Cloud registries. The project lists Envoy release 1.39 as its proxy base.
Is Istio 1.31 safe to upgrade to from 1.30?
It is safe with preparation. The risky items are documented: unhealthy endpoints are now sent to proxies by default, HBONE-tunneled WorkloadEntries need a label, very large meshes may exceed a gRPC message limit, and custom MCP consumers need control-plane identity. Use a revision-based canary, test those four items in staging, and keep the old revision until a full business cycle has passed.
How do I do a canary upgrade of Istio with revisions?
Install the new control plane with istioctl install --set revision=<name>, create revision tags with istioctl tag set, label one test namespace to the canary tag, and restart its deployments. After validation, repoint the stable tag with --overwrite and restart in waves. Rollback means repointing the tag at the old revision and restarting again; retire the old revision with istioctl uninstall --revision.
Does upgrading Istio ambient mode require restarting application pods?
No for the ambient components themselves. Ztunnel and the CNI node agent are DaemonSets and waypoints are deployments, so application pods do not restart for a data plane upgrade. Any sidecar workloads in a mixed mesh still need restarts. Ztunnel upgrades do interrupt connections on the node during rollout, so stagger them. Follow the ambient Helm upgrade guide for exact steps.
Should I use sidecar or ambient mode with Istio 1.31?
Choose per namespace. Keep sidecars where you depend on per-pod Envoy customization such as EnvoyFilter or Wasm, or need strict per-pod proxy isolation. Choose ambient where sidecar overhead and restart coordination are the real costs and you can accept L4 only or add a waypoint for L7. A mixed mesh is supported, so you can migrate gradually and reverse a namespace if needed.
When does Istio 1.29 reach end of life?
According to the Istio supported-releases page, 1.29 reaches end of life on 12 October 2026. Istio 1.30 follows at roughly December 2026, and 1.31 at roughly February 2027. Support means the community produces patch releases for critical issues and provides technical assistance, and a minor is supported until six weeks after the second subsequent minor ships. Treat approximate dates as planning figures and confirm on the page.
Further Reading
- Istio ambient mesh versus Linkerd: a service mesh ADR for the choice of mesh before the choice of version.
- Cilium service mesh versus Istio ambient mesh ADR for the eBPF-based alternative.
- Istio ambient multicluster with the Gateway API Inference Extension for LLM serving for the multicluster pattern 1.31 stabilizes.
- Ambient IoT in 3GPP Release 19 versus Release 20, a different “ambient” entirely: battery-free cellular IoT, unrelated to the mesh mode despite the shared word.
- Announcing Istio 1.31.0, the primary source for the release.
- Istio canary upgrade documentation and Istio supported releases.
By Riju — about
