Kubernetes Graceful Shutdown and PodDisruptionBudgets: Why Rollouts Still Drop Requests

Kubernetes Graceful Shutdown and PodDisruptionBudgets: Why Rollouts Still Drop Requests

Kubernetes Graceful Shutdown and PodDisruptionBudgets: Why Rollouts Still Drop Requests

You run a rolling update on a healthy Deployment with three replicas, readiness probes, and a PodDisruptionBudget, and your load test still shows a handful of 502s and connection resets. Nothing crashed. Every pod passed its probes. The errors line up exactly with the moments old pods were being replaced. This is the most common reliability complaint in production clusters, and it is almost never a bug in your application logic. It is a race between three independent actors that Kubernetes deliberately does not synchronise.

Kubernetes graceful shutdown is the discipline of making those actors agree: the kubelet that signals your container, the control plane that removes the pod from service endpoints, and the proxies and load balancers that stop routing to it. The first two run in parallel, and the third lags behind both. Add a PodDisruptionBudget (PDB) that only protects against voluntary evictions, a grace period that silently counts down while you wait, and long-lived connections that no budget covers, and “zero downtime” becomes something you have to engineer rather than assume.

This article walks the real termination sequence, shows where requests are lost, and gives you tested manifests and signal handlers to close each gap.

What this covers: the pod termination timeline, the endpoint propagation race, preStop sleep and the native sleep action, grace-period arithmetic, SIGTERM handlers in Go and Python, HTTP keep-alive and gRPC GOAWAY draining, PDB semantics and the eviction API, node drains with autoscalers, rollout surge settings, load balancer deregistration, a load-test method, and a decision matrix.

Context and Background

Kubernetes treats pod deletion as a request, not a command. When something deletes a pod, whether a controller scaling down, a rolling update replacing it, or kubectl drain evicting it, the API server does not remove the object immediately. It sets a deletionTimestamp and a grace period, and the pod enters the Terminating state. From that instant, several components observe the change independently through watches, and each acts on its own schedule. The Kubernetes documentation on pod lifecycle describes this termination flow, and the default grace period is 30 seconds unless you override terminationGracePeriodSeconds.

The design is intentional. Kubernetes is a level-triggered, eventually consistent system. There is no global barrier that says “traffic has stopped, now deliver SIGTERM”. The endpoints controller, the kubelet, kube-proxy (or an eBPF replacement), ingress controllers, service meshes, and cloud load balancers all converge on the new state independently. In steady state that convergence takes under a second in a small cluster. In a large cluster, under load, with an external load balancer in the path, it can take many seconds. Those seconds are your outage window.

Two other mechanisms often get conflated with graceful shutdown, and separating them is the first step to reasoning clearly. A rolling update strategy (maxSurge and maxUnavailable) controls how many pods the Deployment controller replaces at once. A PodDisruptionBudget controls how many pods voluntary disruptions, such as node drains, may take down at once. Neither one controls what happens to the in-flight requests of a pod that has already been chosen for termination. That is the job of shutdown choreography, and it lives partly in your manifest and partly in your application code.

If you run autoscalers, the problem multiplies because pods are terminated far more often. Node consolidation in Karpenter evicts pods through the same eviction API a human would use, and event-driven scale-in from KEDA deletes pods through the replica controller. A service that tolerates one rollout a day without visible errors may fail constantly once a consolidating autoscaler is cycling nodes every few minutes. Graceful shutdown stops being a polish item and becomes a prerequisite for using autoscaling at all.

A note on scope and evidence. Where I cite a version number or a default below, it comes from the upstream documentation or the Kubernetes Enhancement Proposal (KEP) for that feature. Where a behaviour depends on your CNI, ingress controller, mesh, or cloud, I say so rather than assert a universal number. The specific delays in the worked examples are illustrative values chosen to show the arithmetic, not measurements from any particular cluster.

The Termination Sequence and the Race That Drops Requests

Direct answer: When a pod is deleted, the API server marks it Terminating. In parallel, the kubelet runs the preStop hook and then sends SIGTERM, while the endpoints controller removes the pod from EndpointSlices. Proxies and load balancers learn of the removal later, so traffic keeps arriving after SIGTERM unless you delay shutdown and keep serving until routing has converged.

Kubernetes graceful shutdown timeline showing API delete, parallel kubelet SIGTERM and endpoint removal, and lagging proxy updates

Figure 1: The pod termination timeline. The kubelet path and the endpoint removal path start together, but proxy and load balancer convergence finishes last.

Figure 1 shows the sequence. A client or controller issues a DELETE for the pod. The API server records the deletionTimestamp and returns. Two things now happen at the same time, and that simultaneity is the heart of the problem.

On the node, the kubelet sees the pod is terminating. It starts the grace period countdown, runs each container’s preStop hook if one is defined, and then asks the container runtime to send SIGTERM to the main process of each container. If the container is still alive when the grace period ends, the runtime sends SIGKILL. The grace period timer starts at deletion, not after the preStop hook finishes, so time spent in preStop is time you do not get back for the application’s own shutdown.

In the control plane, the EndpointSlice controller observes the same pod update. Because the pod is terminating, it stops listing the pod as ready for new traffic. In modern clusters the EndpointSlice API records serving and terminating conditions, so consumers can tell a pod that is draining but still able to serve from one that is gone. Then every consumer of EndpointSlices has to notice the change: kube-proxy on every node rewrites iptables or IPVS rules, your ingress controller updates its upstream list, a service mesh control plane pushes new configuration to sidecars, and an external load balancer controller calls a cloud API to deregister a target.

Why the race is real

Each hop in that chain is an asynchronous watch followed by some work. kube-proxy batches rule syncs. An ingress controller such as NGINX may reload its configuration or update upstreams dynamically. A cloud load balancer integration queues a deregistration and the cloud then applies it with its own delay. None of these consumers can block the kubelet, because the kubelet does not wait for anyone. So the timeline in the first second after deletion has your application receiving SIGTERM while a fraction of the fleet still routes new connections to it.

An application that treats SIGTERM as “exit now” will close its listener and terminate. Every request that arrives in the next several seconds hits a closed port, which the client sees as a connection refused or reset, or which the proxy converts into a 502 or 503. The failure rate is small, perhaps a fraction of a percent of requests per replaced pod, but it repeats for every pod in every rollout and every drain. Over a day of autoscaler activity that adds up to a visible error budget burn with no corresponding incident.

What the fix has to accomplish

The fix is not to speed up the control plane. It is to make termination tolerant of the lag. The pod must keep accepting and serving requests after it learns it is terminating, for at least as long as it takes routing to converge. Then it must stop accepting new connections, finish in-flight requests, close idle keep-alive connections, and exit, all before the grace period expires. That yields three requirements that the rest of this article builds on: a delay before shutdown starts, a drain phase that respects in-flight work, and a grace period big enough to contain both.

There is a subtle point about readiness. A pod that fails readiness is removed from endpoints. A terminating pod is removed from endpoints too. But if your application flips its own readiness to failing on SIGTERM, you add a second removal signal that arrives after the first, which helps nothing, because the endpoint removal triggered by deletion is already in flight. The extra readiness flip is useful only for pods that begin draining without being deleted, for example when your application decides to restart itself. Do not mistake it for the cure.

Building the Shutdown Sequence: preStop, SIGTERM, and the Grace Period Budget

Direct answer: Add a preStop sleep of a few seconds so the pod keeps serving while endpoints propagate, then handle SIGTERM by stopping the listener gracefully and draining in-flight requests. Set terminationGracePeriodSeconds to at least the preStop sleep plus your worst-case drain time plus a safety margin.

Kubernetes graceful shutdown budget showing preStop sleep, drain window and SIGKILL deadline inside terminationGracePeriodSeconds

Figure 2: The grace period is one budget shared by the preStop hook and the application drain. When it reaches zero, SIGKILL ends the process.

The preStop sleep and why it works

A preStop hook runs inside the container before SIGTERM is sent. Because it must finish before the signal, a hook that simply waits delays the signal. During that wait the application is untouched: it still accepts connections and serves requests, while the endpoint removal propagates through the cluster. When the hook returns, SIGTERM arrives, and by then almost all routers have stopped sending new traffic. The trick is crude and effective. It converts a race into a bounded delay.

Historically teams implemented the wait with an exec hook running sleep 10. That has a hidden requirement: the container image must contain a sleep binary. Distroless and scratch-based images do not, so the hook fails, the failure is logged as a FailedPreStopHook event, and the container proceeds to SIGTERM with no delay. People discover this only when the errors come back after a base-image change.

The native sleep action removes the dependency. KEP-3960 added a sleep handler to lifecycle hooks, and the KEP lists it as alpha in v1.29, beta in v1.30, and stable in v1.34. A follow-up, KEP-4818, allows a zero-second sleep, listed as alpha in v1.32, beta in v1.33, and stable in v1.34. That second change matters mostly to admission webhooks that inject a hook and want a harmless no-op value. On any cluster at v1.30 or later you can use the native action, but check your own version and feature-gate state before relying on it, because managed providers lag upstream.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: api
spec:
  replicas: 3
  selector:
    matchLabels: { app: api }
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  template:
    metadata:
      labels: { app: api }
    spec:
      terminationGracePeriodSeconds: 60
      containers:
        - name: api
          image: registry.example.com/api:1.8.2
          ports:
            - containerPort: 8080
          readinessProbe:
            httpGet: { path: /readyz, port: 8080 }
            periodSeconds: 2
            failureThreshold: 2
          lifecycle:
            preStop:
              sleep:
                seconds: 15

In this manifest the budget arithmetic is simple. The pod gets 60 seconds in total. The preStop sleep consumes 15. The application then has up to 45 seconds to drain after SIGTERM, minus whatever margin you want. If a request can legitimately run for 30 seconds, you have 15 seconds of slack. If requests can run for 60 seconds, the manifest is wrong and the pod will be killed mid-request no matter how well the code behaves.

Grace period arithmetic

Write the inequality down and keep it in the repository next to the manifest. The required grace period is at least the preStop duration, plus the longest request you promise to finish, plus the time to flush buffers and close dependencies, plus a margin. Sixty seconds is a plausible value for a typical API with 10 to 20 second requests. A batch consumer that must finish a unit of work, or a stateful database that must flush to disk, needs a different number, sometimes minutes.

Two traps hide here. First, the grace period is a per-pod value, but evictions and node drains can also supply their own. The eviction API accepts a grace period in the delete options, and kubelet shutdown on node termination has its own limits, so the number in your manifest is a request that other parties can shorten. Second, a very long grace period is not free. It delays rollouts, keeps terminating pods holding capacity and IP addresses, and makes kubectl drain slow. Size it to the worst case you have measured, not to a round number.

Why you still need the application to handle SIGTERM

The sleep protects against the routing lag. It does nothing for in-flight requests, and it does nothing if your process ignores the signal. Many containers run the application behind a shell wrapper, so the shell is PID 1 and the application never receives SIGTERM. A shell script entrypoint that does not use exec is the classic cause: the signal goes to the shell, which does not forward it, and the pod sits until SIGKILL at the end of the grace period. Rollouts then take the full grace period per pod, and the symptom reads as “deployments are slow” rather than “signals are lost”. Use the exec form of CMD or ENTRYPOINT, use exec in wrapper scripts, or run a tiny init such as tini that forwards signals and reaps zombies.

A subtle point about PID 1 on Linux: the kernel does not apply default signal actions to PID 1 in a container’s PID namespace. A process running as PID 1 that has not installed a SIGTERM handler will ignore SIGTERM rather than terminate. This is why runtimes that work when launched from a terminal appear to hang in a container. Install the handler explicitly.

Application Side: Draining Connections in Go and Python

Direct answer: On SIGTERM, stop accepting new connections, let in-flight requests finish within a deadline, tell clients to reconnect elsewhere by sending Connection: close or an HTTP/2 GOAWAY frame, then exit. Go’s http.Server.Shutdown and gRPC’s GracefulStop implement most of this; Python needs explicit wiring.

SIGTERM handling flow for graceful shutdown showing stop accepting, drain in-flight requests, close idle keep-alives and exit

Figure 3: The application-side drain loop. Listener closes first, then in-flight work completes under a deadline, then connections close and the process exits.

The preStop hook buys time for routing; the application’s signal handler uses that time well. A correct handler does five things in order. It marks the process as draining so health endpoints can reflect it. It stops accepting new connections. It lets in-flight requests finish under a deadline shorter than the grace period remaining. It closes idle persistent connections and tells active ones to go away. Finally it releases dependencies and exits with status zero.

A Go handler

Go’s standard library does most of the work. http.Server.Shutdown closes listeners, closes idle connections, and waits for active connections to become idle, honouring a context deadline. It does not interrupt a handler that is mid-request, which is what you want.

package main

import (
    "context"
    "log"
    "net/http"
    "os"
    "os/signal"
    "sync/atomic"
    "syscall"
    "time"
)

func main() {
    var draining atomic.Bool

    mux := http.NewServeMux()
    mux.HandleFunc("/readyz", func(w http.ResponseWriter, r *http.Request) {
        if draining.Load() {
            http.Error(w, "draining", http.StatusServiceUnavailable)
            return
        }
        w.WriteHeader(http.StatusOK)
    })
    mux.HandleFunc("/work", func(w http.ResponseWriter, r *http.Request) {
        if draining.Load() {
            // ask the client to open a fresh connection elsewhere
            w.Header().Set("Connection", "close")
        }
        time.Sleep(2 * time.Second) // simulated work
        w.Write([]byte("ok\n"))
    })

    srv := &http.Server{Addr: ":8080", Handler: mux}

    go func() {
        if err := srv.ListenAndServe(); err != http.ErrServerClosed {
            log.Fatalf("listen: %v", err)
        }
    }()

    stop := make(chan os.Signal, 1)
    signal.Notify(stop, syscall.SIGTERM, syscall.SIGINT)
    <-stop

    draining.Store(true)
    // preStop already delayed us; keep a short pause for stragglers
    time.Sleep(2 * time.Second)

    ctx, cancel := context.WithTimeout(context.Background(), 40*time.Second)
    defer cancel()
    if err := srv.Shutdown(ctx); err != nil {
        log.Printf("forced shutdown: %v", err)
    }
    log.Println("clean exit")
}

The 40-second shutdown deadline sits inside the 45 seconds that remain after the 15-second preStop sleep in the earlier manifest, leaving a margin for the two-second pause. If the deadline expires, Shutdown returns and the process exits anyway, which is better than being killed uncontrolled because the log line tells you the drain was too slow.

A Python handler

Python web servers vary. With a WSGI server such as Gunicorn, the master process handles SIGTERM by default for a graceful stop and gives workers up to the configured graceful_timeout to finish. With ASGI and Uvicorn, SIGTERM triggers a graceful shutdown that waits for in-flight requests, and you can bound it with --timeout-graceful-shutdown. Check the options for the version you run, since they have changed over time. For a hand-rolled asyncio service, wire the signal explicitly.

import asyncio
import signal
from aiohttp import web

draining = False

async def readyz(request):
    if draining:
        return web.Response(status=503, text="draining")
    return web.Response(text="ok")

async def work(request):
    await asyncio.sleep(2)
    headers = {"Connection": "close"} if draining else {}
    return web.Response(text="ok\n", headers=headers)

async def main():
    global draining
    app = web.Application()
    app.add_routes([web.get("/readyz", readyz), web.get("/work", work)])
    runner = web.AppRunner(app, shutdown_timeout=40)
    await runner.setup()
    site = web.TCPSite(runner, "0.0.0.0", 8080)
    await site.start()

    stop = asyncio.Event()
    loop = asyncio.get_running_loop()
    for sig in (signal.SIGTERM, signal.SIGINT):
        loop.add_signal_handler(sig, stop.set)

    await stop.wait()
    draining = True
    await asyncio.sleep(2)       # let late-routed requests land
    await runner.cleanup()       # stops listener, waits up to shutdown_timeout

asyncio.run(main())

The pattern is identical to the Go version: flip a flag, wait briefly, close the listener, wait for in-flight work under a deadline. Treat the specific aiohttp parameter names as something to confirm against the version you pin, because I have not verified every release.

Keep-alive connections and HTTP/2

A pod is not drained until its connections are. HTTP/1.1 clients reuse connections, and a client that holds an open keep-alive connection to the old pod will keep sending requests on it even after the pod leaves the endpoint list, because the connection was already established. Removing a pod from EndpointSlices stops new connections only. Setting Connection: close on responses during the drain tells the client to open a new connection, which then lands on a live pod. Go’s Shutdown also closes idle connections for you.

HTTP/2 and gRPC multiplex many requests over one long-lived connection, which makes the problem sharper. The protocol has a graceful mechanism: the server sends a GOAWAY frame announcing the last stream it will process, the client stops opening new streams on that connection, and in-flight streams complete. In gRPC servers this is what GracefulStop does. RFC 9113, the HTTP/2 specification, defines GOAWAY, and the gRPC documentation covers graceful shutdown behaviour for each language. An application that exits abruptly instead sends nothing, and the client discovers the failure only on the next request, which is the source of “unavailable” errors during rollouts of gRPC services.

Client-side load balancing complicates this. A gRPC client that resolved the Service’s ClusterIP sees one virtual address, so all of its connections are pinned to whichever pods it first connected to, and new pods get no traffic. Headless services with client-side round robin, or a mesh that balances per request, change the behaviour. Whichever you choose, make sure the client reconnects on GOAWAY and re-resolves, otherwise drained pods leave and the remaining pods take a lopsided share.

Stateful and queue workloads

Not every workload is a request server. A queue consumer must stop fetching, finish or hand back the message in progress, and acknowledge or release it, all before exit. A StatefulSet running a database must flush and release leadership. In both cases the handler is the same shape as above with different dependencies. The key is that the unit of work has a bounded duration, so you can write the grace period inequality for it. If a unit of work can run for an hour, shutdown is a checkpointing problem, not a draining problem, and you should design the job to resume rather than to finish.

PodDisruptionBudgets, Evictions, and Node Drains

Direct answer: A PodDisruptionBudget limits how many pods of an application may be down at once during voluntary disruptions such as node drains. It is enforced by the eviction API, not by Deployment rollouts or direct pod deletes, and it says nothing about whether the evicted pod shuts down gracefully.

PodDisruptionBudget eviction flow showing drain request, eviction API budget check, 429 retry and graceful pod termination

Figure 4: How a node drain interacts with a PodDisruptionBudget. The eviction API checks the budget, refuses with a retryable error when it would be violated, and otherwise starts normal graceful termination.

What a PDB actually guards

A PDB selects pods by label and states either minAvailable or maxUnavailable, each as an integer or a percentage. You set one of the two, not both. The Kubernetes task documentation notes that the feature has been stable since v1.21 and that your cluster owner must confirm their tooling respects PDBs. The key phrase is voluntary disruption. Node drains, cluster autoscaler scale-down, and Karpenter consolidation go through the eviction API and honour the budget. Involuntary events such as a node crash, a kernel panic, or an out-of-memory kill do not consult it, though they do count against it when the controller computes how many pods are currently healthy.

The eviction API is a subresource of the pod. When a client posts an Eviction, the API server checks every PDB that selects the pod. If the eviction would leave fewer healthy pods than the budget allows, the request is refused with HTTP 429 and the client is expected to retry. If the budget allows it, the server proceeds to delete the pod with normal graceful semantics, honouring the pod’s grace period. Kubernetes documents API-initiated eviction as respecting both PDBs and terminationGracePeriodSeconds. From general API behaviour, a pod matched by multiple PDBs yields a server error rather than a decision, so give each pod exactly one PDB. kubectl drain uses the eviction API and keeps retrying refused evictions until its timeout, and it offers a flag to bypass eviction by deleting pods directly, which also bypasses the budget. Treat that flag as an emergency tool.

What a PDB does not guard

The most common misunderstanding is that a PDB protects a rolling update. It does not. Deployment rollouts delete old pods through the ReplicaSet controller, which does not use the eviction API, so the rollout pace is governed by maxSurge and maxUnavailable on the strategy. A PDB with minAvailable: 2 on a three-replica Deployment will not slow a rolling update at all. It will, however, make a concurrent node drain wait, so the two mechanisms interact only when both happen at once.

The second misunderstanding is that a PDB guarantees capacity. It guarantees a count of healthy pods at the moment of eviction. If those pods are healthy but overloaded, or if the remaining pods were all scheduled on the node that fails next, the budget is satisfied and users still suffer. Pair the PDB with topology spread constraints so that replicas live on different nodes and zones, otherwise a single drain or zone event removes more capacity than the budget implies.

Choosing minAvailable or maxUnavailable

Prefer maxUnavailable for workloads that scale with a HorizontalPodAutoscaler or KEDA. A percentage minAvailable recomputes against the current replica count, but an integer minAvailable does not, and a fixed integer can accidentally equal your replica count when the autoscaler scales down, which blocks every eviction. That is the well-known “PDB deadlock” where a drain hangs forever because minAvailable equals the number of replicas. Setting maxUnavailable: 1 expresses “one at a time” regardless of scale and avoids it. For a singleton, a PDB with maxUnavailable: 0 blocks all voluntary eviction, which is occasionally what you want for a job mid-run and is usually a mistake for a long-lived service.

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api
spec:
  maxUnavailable: 1
  unhealthyPodEvictionPolicy: AlwaysAllow
  selector:
    matchLabels:
      app: api

unhealthyPodEvictionPolicy

Without extra configuration, a PDB can trap pods that are already broken. Consider a pod stuck in CrashLoopBackOff: it is running but not ready, so it counts as unhealthy, and with the default policy the eviction of a running-but-not-ready pod is allowed only when the application is not already disrupted. If the application is disrupted, the unhealthy pod cannot be evicted, so the drain stalls on exactly the pod you most want to remove.

KEP-3017 introduced the unhealthyPodEvictionPolicy field to resolve this. The KEP lists it as alpha in v1.26, beta in v1.27, and stable in v1.31. It accepts two values. IfHealthyBudget is the default when unset: running-but-not-ready pods can be evicted only when the current healthy count is at least the desired healthy count. AlwaysAllow: running-but-not-ready pods can be evicted regardless, while healthy pods remain subject to the budget. The KEP also notes that pods of an already disrupted application may never get a chance to become Ready, which is the argument for AlwaysAllow on most stateless services. For stateful systems where an unready pod might still hold the only copy of data, think carefully before choosing it.

Node drains, autoscalers, and the grace period

Cluster autoscalers and Karpenter drain nodes by evicting pods, so your PDB and your shutdown choreography are exactly what stand between consolidation and an outage. Karpenter’s disruption model evaluates PDBs before it voluntarily disrupts a node, and a PDB that never allows eviction will block consolidation entirely, which shows up as nodes that never go away. Karpenter also has a node-level termination grace period setting that caps how long a drain may take before remaining pods are forcibly deleted; check the version you run, as I have only confirmed the general mechanism, not every field name. The trade-off is explicit: a cap bounds how long a stuck drain delays node replacement, but it overrides the application’s own grace period for the pods it cuts short.

Spot instance reclamation is a different path altogether. A cloud provider’s interruption notice gives a short, fixed warning, often around two minutes on AWS for spot instances, and the node then disappears. A 5-minute grace period on a pod is meaningless there: the node-level shutdown time is the real cap. Design spot-backed workloads for a shutdown budget of well under a minute, and use the interruption handler or the autoscaler’s native handling to cordon and drain as soon as the notice arrives.

Rollout Settings, Readiness, and Load Balancers

Direct answer: For zero-downtime Deployment rollouts, set maxUnavailable to 0 with a positive maxSurge, make readiness probes reflect true serving ability, add a preStop sleep, and align any external load balancer’s deregistration delay with the pod’s termination budget.

maxSurge and maxUnavailable

With the default rolling update strategy, a Deployment may exceed its replica count by maxSurge and may drop below it by maxUnavailable, and both default to 25 percent. For a service that must not lose capacity, maxUnavailable: 0 with maxSurge: 1 (or a percentage) forces the controller to start a new pod and wait for it to become Ready before terminating an old one. The cost is a transient extra pod, which means headroom in the cluster and in any per-pod licence or quota. If the cluster has no capacity, the new pod stays Pending and the rollout stalls, which is safe but slow.

Readiness is the gate in that handshake, so a readiness probe that lies defeats the whole scheme. A probe that returns 200 when the process is merely listening, before caches are warm or connection pools are filled, lets traffic arrive too early. A probe that depends on a downstream service can cascade, because one slow dependency marks every replica unready and removes all of them at once. Make readiness describe “this pod can serve its own requests right now”, and use minReadySeconds to require stability before the controller counts a pod as available. Startup probes protect slow-starting applications from being killed by liveness checks during boot.

Readiness gates add one more condition. A pod may declare readinessGates that require an external controller to set a condition to True before the pod is considered ready. Cloud load balancer controllers use this to make Kubernetes wait until the load balancer has actually registered the new pod as healthy, closing the gap on the scale-up side. Without it, the Deployment controller may proceed to terminate the old pod after the new pod is Ready by Kubernetes’ definition but before the load balancer has started sending it traffic, which can briefly leave the service short of capacity.

The external load balancer leg

When traffic enters through a cloud load balancer, there is a second deregistration to account for. An AWS Application Load Balancer, for example, has a target group attribute called deregistration delay with a default of 300 seconds, per the AWS documentation. When a target is deregistering, the load balancer stops sending it new requests, and the target is held in a draining state so in-flight requests can finish. The documentation also warns that if a deregistering target closes the connection before the delay elapses, the client receives a 500-level error. The delay does not reduce propagation lag; it governs how long the load balancer waits after it has already decided to stop routing.

So there are two clocks. The load balancer learns that the pod is going away through the controller’s call, which takes however long the controller and cloud API take, and only then starts its draining. A pod that exits at SIGTERM plus 10 seconds, while the load balancer is still routing to it, produces the 500-level errors the documentation describes. The remedy is the same as inside the cluster: keep serving through a preStop delay that exceeds the registration propagation time. A very long deregistration delay is not the fix and has its own cost, because the controller or a rolling deploy may wait on it. Many teams lower the default from 300 seconds to something near their longest request, such as 30 to 60 seconds. If you use IP targets with a cloud load balancer controller, consult its documentation for pod readiness gate support and how it handles termination.

Long-lived connections

Neither the PDB nor the preStop sleep helps a WebSocket, a server-sent event stream, a gRPC streaming call, or a database session that lives for hours. These connections end only when the server or client closes them, so a rollout cuts them off at the end of the grace period. You have three honest options. Make clients reconnect transparently with jittered backoff, so a cut connection is a blip. Have the server send an application-level “reconnect elsewhere” message during the drain, which spreads reconnects across the grace window instead of at the end. Or accept the cut and size the grace period to the maximum acceptable session length. The first two scale; the third rarely does. Whichever you pick, add jitter to the reconnect, otherwise every client drops and reconnects at the same moment and the surviving pods absorb a thundering herd.

Progressive delivery as a complement

Shutdown choreography fixes the pod-level race. It does not tell you whether the new version is healthy. Progressive delivery controllers such as Argo Rollouts shift traffic gradually and abort on bad metrics, and they have their own interaction with shutdown timing: the controller scales down old pods only after traffic has moved, and you still need the same preStop and drain handling in the old pods. The posts on Argo Rollouts for edge fleets and the Argo Rollouts versus Flagger decision record cover that layer. Treat the two as complementary: canary analysis catches bad releases, and graceful shutdown ensures that even good releases do not drop requests.

Testing It: A Load Test That Runs During the Rollout

You cannot see these errors by reading manifests. They appear only under traffic, so prove the fix with a test that mimics the failure: sustained load, a rollout in the middle, and a count of non-success responses afterwards.

The method has four steps. First, deploy the service with two or more replicas and a stable client path, ideally through the same ingress or load balancer production uses, since a test that bypasses it skips the slowest hop. Second, start a constant-rate load generator at a rate high enough that every pod receives requests every few hundred milliseconds. Tools such as k6, vegeta, or hey work, and a closed-loop tool that reuses keep-alive connections is more representative than one that opens a connection per request, because keep-alive is where drain bugs live. Third, while the load runs, trigger the disruption you care about: kubectl rollout restart deployment/api, then separately kubectl drain of a node hosting the pods. Fourth, count non-2xx responses and connection errors, and look at their timestamps relative to the pod deletion events.

# constant 200 req/s for 3 minutes against the real entry point
echo "GET https://api.example.com/work" | \
  vegeta attack -rate=200 -duration=180s -keepalive | \
  tee results.bin | vegeta report

# in another terminal, while it runs
kubectl rollout restart deployment/api
kubectl rollout status deployment/api

Interpret the results by pattern. A burst of errors within a second or two of pod termination points to the routing race, so lengthen the preStop sleep. Errors a few seconds later, clustered at the end of the grace period, point to connections being killed by SIGKILL, meaning the drain is too slow or SIGTERM was never delivered. Errors only on gRPC or WebSocket traffic point to missing GOAWAY or reconnect handling. Errors that appear only during drains, never during rollouts, point to PDB or topology problems. Re-run after each change, one variable at a time, and keep the load test in CI against a staging cluster, because shutdown regressions arrive with base image changes and framework upgrades, not with feature code.

Be honest about the limits of the test. A clean run on a quiet staging cluster proves less than a clean run in production-like conditions, since propagation delay grows with cluster size, node count, and control plane load. Treat the preStop value as a number you tune upward from evidence, and leave headroom beyond what your test needed.

Trade-offs, Gotchas, and What Goes Wrong

Every mechanism above has a price, and most production incidents come from using one without understanding the cost.

The sleep is a heuristic, not a guarantee. A preStop sleep assumes routing converges within N seconds. Nothing enforces that. Under a loaded control plane or a slow cloud API, propagation can exceed your value and the errors return, but only occasionally, which is the worst way for them to return. Monitor the error rate during rollouts as a standing metric rather than treating the fix as done.

Longer sleeps slow everything down. A 15-second sleep adds 15 seconds to every pod replacement. On a 200-replica Deployment with maxSurge of 10 percent, that stretches a rollout by minutes. It also extends node drains and delays Karpenter consolidation. Pick the smallest value that your measurements support and make it a named constant rather than a guess.

SIGKILL is silent. When the grace period expires, the process vanishes without logs. If your drain regularly takes longer than the budget, you lose the work and learn nothing. Log the drain start and end, emit a metric for how long drain took, and alert when it approaches the grace period. A rising trend predicts the first dropped requests before they happen.

PID 1 and signal forwarding. Shell wrappers, npm start, java -jar behind a script, and some process managers swallow SIGTERM. Verify in a running container that the application, not a shell, is PID 1 or is the target of forwarded signals. A pod that takes exactly terminationGracePeriodSeconds to disappear is the diagnostic signature.

Sidecars complicate ordering. A service mesh proxy sidecar can terminate before the application finishes, cutting off its outbound calls, or can keep running after the application stops. Native sidecar containers, which Kubernetes handles with a defined shutdown order, address part of this; the pod lifecycle documentation has a section on shutdown with sidecar containers that you should read for your version. Without them, you may need the application’s preStop to wait for the proxy’s own drain, or to configure the mesh to hold the proxy until connections end.

PDBs can block the platform. A PDB that allows zero disruptions, an integer minAvailable equal to replicas, or an application with a single replica and minAvailable: 1, will stall every node upgrade. Managed Kubernetes upgrades surface this as an upgrade that hangs or times out. Audit PDBs for ones that allow no eviction, and prefer budgets expressed as maxUnavailable.

Do not paper over a design problem. If a request takes ten minutes, no grace period makes a rollout comfortable. Move the work to a queue or a job with checkpoints. Likewise, if your clients cannot tolerate a reconnect, no amount of preStop tuning will save them. Fix the protocol contract, not the manifest.

Practical Recommendations

Start with the highest-leverage changes and work outward. Most services reach acceptable behaviour with four settings and one code change, and the rest is tuning from measurement.

Make the application receive and act on SIGTERM. Verify PID 1, use the exec form, and implement a drain with a deadline. Add a preStop delay using the native sleep action on clusters where it is available, or an exec sleep only if the image provably contains the binary, and start with a value in the range of 5 to 15 seconds, then tune it from load-test evidence. Set terminationGracePeriodSeconds from the inequality, not from habit. Use maxUnavailable: 0 with a surge for services that cannot lose capacity.

Then add the supporting controls. Give each critical service one PDB expressed as maxUnavailable, with unhealthyPodEvictionPolicy: AlwaysAllow for stateless workloads on clusters where the field is stable. Spread replicas across nodes and zones. Reconcile the external load balancer’s deregistration delay with the pod’s budget. For long-lived connections, ship client reconnect with jitter.

A short checklist you can paste into a review template:

  • Is the application, not a shell, PID 1, and does it handle SIGTERM?
  • Does the drain have a deadline that fits inside the grace period minus the preStop?
  • Is there a preStop delay that exceeds measured endpoint propagation, including the load balancer leg?
  • Is terminationGracePeriodSeconds at least preStop plus worst-case request plus margin?
  • Do HTTP/1.1 responses send Connection: close and gRPC servers send GOAWAY during drain?
  • Is there exactly one PDB per application, expressed so it cannot deadlock at low replica counts?
  • Are replicas spread across nodes and zones?
  • Does a load test during a rollout and a drain show zero errors?

If you have limited time, the load test is the highest-value item, because it turns every other setting from belief into evidence.

Frequently Asked Questions

Why do pods still drop requests during a rolling update even with readiness probes?

Readiness probes control when a pod starts receiving traffic, not how traffic stops. At termination, the kubelet sends SIGTERM while endpoint removal propagates through kube-proxy, ingress controllers, and load balancers in parallel. For a few seconds new requests still reach a pod that is already shutting down. A preStop sleep that keeps the pod serving, combined with a proper SIGTERM drain, closes that gap.

How long should a preStop sleep be?

Long enough to exceed the time your cluster needs to propagate endpoint removal, including any external load balancer, and short enough not to slow rollouts unreasonably. Values between 5 and 15 seconds are common starting points, but the right number is cluster specific. Measure it with a load test during a rollout, increase it until errors vanish, then add headroom for control plane load.

Does a PodDisruptionBudget protect rolling updates?

No. A PDB is enforced by the eviction API, which node drains and autoscalers use. Deployment rolling updates delete pods through the ReplicaSet controller and are paced by maxSurge and maxUnavailable. A PDB will make a concurrent node drain wait, but it does not slow or constrain a rollout, and it does not make any pod shut down gracefully.

What is the native sleep action and which Kubernetes version has it?

It is a lifecycle hook handler that waits a set number of seconds without needing a sleep binary in the image. According to KEP-3960 it was alpha in v1.29, beta in v1.30, and stable in v1.34. A related KEP, 4818, allows a zero-second value, also reaching stable in v1.34. Confirm availability on your managed provider’s version.

What does unhealthyPodEvictionPolicy AlwaysAllow do?

It lets the eviction API remove pods that are running but not ready even when the budget would otherwise block it, while healthy pods stay protected. This prevents drains from hanging on crash-looping or never-ready pods. The field reached stable in v1.31 per KEP-3017. The default, IfHealthyBudget, allows such evictions only while the application is not already disrupted.

How do I gracefully shut down gRPC services on Kubernetes?

Call the server’s graceful stop on SIGTERM, which sends a GOAWAY frame so clients stop opening new streams while existing ones finish, and bound the wait with a deadline inside the grace period. Make sure clients reconnect and re-resolve on GOAWAY. If clients pin to a ClusterIP, use a headless service or a mesh that balances per request.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *