etcd 3.7 Explained: Raft Consensus, Kubernetes Control Plane Impact and Upgrade Guide
Every object in a Kubernetes cluster, from a Pod to a Secret to a Lease, lives in one small replicated database, and when that database stalls the whole control plane stalls with it. etcd 3.7 shipped on 8 July 2026 and is, on paper, a quiet release: no new consistency model, no new storage engine. In practice it changes three things that platform teams feel directly: how large list reads are served, how cheap lease and key-only operations are, and how much legacy code still sits between you and a clean upgrade.
The current patch line is v3.7.2, tagged on 22 September 2026, and it is the version most teams should be reading release notes for. This post walks through how etcd actually works (Raft, MVCC, bbolt, compaction, fsync sensitivity), what 3.7 changed, how the new RangeStream API interacts with the Kubernetes API server, and how to run a safe upgrade with backup, defrag, and restore drills.
What this covers: the Raft and storage mechanics that explain every etcd failure mode, the verified 3.7 change list and its Kubernetes impact, a rolling upgrade runbook with commands, capacity and quorum sizing, and the traps that cause outages.
Context and Background
etcd is a strongly consistent, distributed key-value store built on the Raft consensus algorithm. It is the only stateful component in the Kubernetes control plane: the kube-apiserver is stateless, and the scheduler and controllers hold their state in API objects, which means in etcd. Lose etcd quorum and the cluster freezes in place. Running Pods keep running, but nothing can be scheduled, scaled, or healed.
The project follows a roughly annual minor-release cadence with a long stabilization tail. etcd 3.6 landed in 2025 and introduced Kubernetes-style feature gates, a supported downgrade path, livez and readyz health endpoints, and the v3discovery protocol. Per the 3.6 announcement, it cut average memory use by at least 50 percent, largely by lowering the default --snapshot-count from 100,000 to 10,000, and reported roughly 10 percent average throughput improvement. etcd 3.7 builds on that base and finishes work that 3.6 started, chiefly the removal of the v2 store.
Support matters for planning. The July 2026 patch announcement states that the v3.4 branch has reached end of support and receives no further patches. The 3.5 and 3.6 branches are still being patched: on 22 September 2026 the project tagged v3.5.34, v3.6.15 and v3.7.2 together. If you are still on 3.4, you are running a database with known unpatched CVEs under your control plane.
Two neighbouring topics shape how you run etcd in industrial and edge environments. Network behaviour affects Raft heartbeats, and the choice of CNI matters when members share a fabric with workload traffic; our Cilium vs Calico ADR for industrial Kubernetes networking covers the data-plane trade-offs. And if your clusters host device-facing control planes, the Azure IoT Akri resource interface shows how many extra custom resources an edge cluster can end up storing in etcd.
A caution on sources. The etcd project’s own announcement, changelog and upgrade guide are the authority here, and where they disagree with a secondary summary I follow the changelog. The official 3.7 announcement is the best single starting point.
How etcd 3.7 Works: Raft, MVCC, and bbolt as One Pipeline
A write in etcd passes through three layers in order: the Raft log for agreement, the MVCC layer for versioning, and the bbolt backend for durable storage. Raft guarantees every member applies the same entries in the same order; MVCC turns each entry into a new revision; bbolt persists the result. Nearly every operational problem maps to one of those three layers.

Figure 1: The etcd write path. The leader appends to its write-ahead log and fsyncs, replicates to followers, commits on quorum acknowledgement, and then applies to the MVCC index and bbolt backend.
The diagram shows why one slow disk can hurt the whole cluster. The commit point sits after a quorum of write-ahead-log fsyncs, so the leader’s own disk and at least one follower’s disk are both on the critical path of every write.
Raft consensus: leader, term, quorum
Raft consensus elects one leader per term. The leader accepts all writes, appends them to its log, sends MsgApp messages to followers, and considers an entry committed once a majority of members have persisted it. A cluster of N members needs floor(N/2)+1 for quorum. The etcd FAQ gives the practical consequence: a three-member cluster tolerates one failure with quorum of two, and a five-member cluster tolerates two failures with quorum of three.
This is why even-sized clusters are a mistake. Four members need three for quorum, so they tolerate only one failure, exactly like three members, while paying more replication cost and giving you two ways to split into undecidable halves. Odd sizes are the only sizes that add fault tolerance.
Timeouts matter as much as sizing. Per the tuning guide, the default heartbeat interval is 100 ms and the default election timeout is 1000 ms. The guidance is that the heartbeat should be roughly 0.5 to 1.5 times the round-trip time between members, and the election timeout at least ten times the round-trip time. A follower that does not hear a heartbeat within the election timeout starts an election, increments the term, and, if it wins, becomes leader.
An election is not free. While it runs, no writes commit, and the kube-apiserver sees timeouts or etcdserver: leader changed errors. Frequent elections are therefore a signal about disk or network, not about etcd itself, and they are where most of the “etcd is flaky” reports on Kubernetes originate.
A second Raft detail that surprises people is the role of the applied index. Committed is not the same as applied: a follower can know an entry is committed before it has finished applying it to the MVCC store. Linearizable reads use the ReadIndex mechanism, which asks the leader for its commit index and waits until the local applied index catches up. The 3.7 release bumps the standalone Raft library to v3.7.0 and its notes mention ReadIndex flow improvements, though the changelog does not quantify them.
MVCC: every write creates a revision
etcd’s data model is multi-version. Every successful write transaction increments a global, monotonically increasing revision, and keys are never updated in place. Instead, the store records that key K was modified at revision R. An in-memory structure, the treeIndex (a B-tree), maps each key to its list of revisions, and the actual values live in the bbolt file keyed by revision.
This design gives Kubernetes its core primitives. A resourceVersion on an API object is an etcd revision. A watch starting from a given revision replays history from that point. Optimistic concurrency, where an update succeeds only if the object is unchanged since you read it, is an etcd transaction comparing mod revision.
The price of keeping history is that the database only grows until you compact. Compaction discards revisions older than a chosen one, but keeps the latest version of every key. Attempting to read a compacted revision returns the error etcdserver: mvcc: required revision has been compacted, which is exactly the signal watchers use to know they must relist.
bbolt: a single-writer B+tree file
The backend is bbolt, a fork of BoltDB: a copy-on-write B+tree stored in one memory-mapped file with a single writer and many concurrent readers. etcd batches writes into bbolt transactions and commits them periodically rather than per Raft entry, which is part of why the Raft log and the database file can differ in position at any moment.
Two properties follow. First, pages freed by compaction are not returned to the filesystem; they are kept inside the file for reuse, so the file never shrinks on its own. Second, the database size, not the key count, is what hits the space quota. The 3.7 release bumps bbolt to v1.5.0 per the changelog. The announcement page also mentions database file size limits and statistics tuning in the bbolt dependency; I could not find a changelog line that sets a default, so treat the specifics as undisclosed until you read the bbolt release notes yourself.
What Changed in etcd 3.7: The Verified Change List
etcd 3.7 has four themes: streaming reads, cheaper lease and key-only operations, removal of legacy v2 code, and a modernised dependency tree. The sections below are drawn from the project’s changelog for the 3.7 series and the release announcement. Where a number appears, it is the project’s own; where none is given, I say so.
Release timeline
The release candidate, v3.7.0-rc.0, was published on 1 June 2026. The GA release, v3.7.0, followed on 8 July 2026. Two patch releases have shipped since: v3.7.1 on 23 July 2026 and v3.7.2 on 22 September 2026. The changelog also has a v3.7.3 section marked “TBC”, which currently lists only a deprecation of NewJournalWriter in client/pkg/v3/logutil. A “TBC” section is a placeholder, not a release.
| Version | Date | Notable content |
|---|---|---|
| v3.7.0-rc.0 | 2026-06-01 | Breaking removals, RangeStream, Unix sockets, Go 1.26 |
| v3.7.0 | 2026-07-08 | GA; CRL bypass fix, non-blocking client creation, raft v3.7.0, bbolt v1.5.0 |
| v3.7.1 | 2026-07-23 | Peer lease handler bound, watch authorization fix, TLS handshake timeout |
| v3.7.2 | 2026-09-22 | WAL-based minimal version fix, snapshot dir fsync fix, gRPC and OpenTelemetry CVE bumps |
RangeStream: streaming large reads
The headline feature is RangeStream, a new server-streaming RPC. The unary Range RPC must hold the whole result, the key-value slice, the serialized protobuf, and the gRPC send buffer, in memory at the same time before anything is sent. KEP-5966, the Kubernetes enhancement proposal that consumes it, describes the second problem precisely: paginated reads recompute the total key count on every page, so per-page work becomes proportional to the total number of keys.
RangeStream accepts the existing RangeRequest and returns a stream of RangeStreamResponse chunks. The server paginates internally with adaptive sizing: it starts with a conservative key limit and then doubles or halves it based on how the previous chunk’s size compares with a target derived from MaxRequestBytes. Intermediate chunks carry only key-value pairs; the final chunk carries the header (cluster ID, member ID, Raft term, revision), the count, and the “more” flag. Clients reassemble the response with proto.Merge().
Consistency is preserved by pinning the MVCC revision after the first chunk. The server does not hold one long transaction for the entire stream; it opens short transactions per chunk against the pinned revision. If compaction removes that revision mid-stream, the server returns ErrCompacted, and the client retries, which matches how paginated list calls already behave.
I could not find published benchmark figures for RangeStream. The Kubernetes v1.37 RangeStream blog post says it reduces memory required for large list reads on both the API server and etcd and makes peak usage more predictable, but quotes no percentages in the text I retrieved. The etcd announcement says Kubernetes users should see a significant decrease in overall etcd CPU usage compared with 3.6, again without a figure. Treat any specific percentage you see repeated elsewhere as unverified until you reproduce it on your own data.
Keys-only ranges and lease performance
When a Range request sets keys_only (or you run etcdctl get --keys-only), 3.7 can answer from the in-memory treeIndex without touching bbolt for values. The exception is a request sorted by value, which still needs the values. The practical beneficiaries are controllers and tooling that list key names to count or enumerate objects, such as per-resource count metrics.
Lease handling gets two changes. LeaseRevoke requests are now prioritised so that lease expiry still happens under overload, which matters because Kubernetes Node heartbeats, Events and some leader-election locks depend on timely lease behaviour. A new FastLeaseKeepAlive feature skips the wait for the applied index when renewing a lease. Separately, the changelog credits a change to (*readView) Rev() using SharedBufReadTxMode with up to a 2x improvement for lease and user or role operations. That “up to 2x” is the project’s claim for those specific operations, not a general throughput multiplier.
FastLeaseKeepAlive deserves a caveat. Skipping the applied-index wait trades a small amount of ordering strictness on the renewal path for speed. The changelog describes it as a feature without describing its default state, so check the 3.7 documentation for how it is enabled before assuming it is on.
Bootstrap from v3store and the end of v2
The 3.7 server bootstraps entirely from the v3 store. It stops loading v2 snapshot files and initialises Raft’s configuration state from v3 data. The v2 HTTP API, v2-on-v3 emulation, v2discovery, client/v2, and v2 request handling are removed. The v2 snapshot generation path remains for backward compatibility: --snapshot-count is kept, and the changelog notes that removal of --max-snapshots is deferred to 3.8 rather than 3.7.
The consequence is operational rather than API-level for Kubernetes, since kube-apiserver has used the v3 API for years. But anything in your estate that still hits /v2/keys, including old health scripts or discovery tooling, breaks. The 3.6.13 and 3.5.32 patch releases added a write-only-skip-check value for --v2-deprecation and strengthened etcdutl check v2store to scan both snapshot and WAL records, precisely so you can find v2 data before you cut over.
Unix sockets, auth, and tooling
Other items: Unix socket endpoints for local communication without a TCP port; the ability for clients to set a JWT directly; retrieval of AuthStatus without authentication; a timeout flag on etcdutl commands that wait for the database file lock; reorganised etcdctl commands and hidden global flags so --help output resembles kubectl. New metrics include etcd_server_request_duration_seconds and four etcd_debugging_server_watch_send_loop_* histograms, which let you see where watch fan-out time goes.
Dependencies and breaking client behaviour
Binaries are built with Go 1.26. The 3.7.0 changelog says Go 1.26.5, while the announcement text I retrieved says 1.26.4; I follow the changelog. v3.7.2 builds with Go 1.26.8. The protobuf stack moved from gogo/protobuf and golang/protobuf to google.golang.org/protobuf, and gRPC middleware moved from v1 to v2 interceptors. The raft library is v3.7.0.
Two changes break embedders. Anyone importing the etcd Go modules, meaning the client SDK, api/, or pkg/ packages, may need code changes because of the protobuf migration. And clientv3 creation is now non-blocking: etcd no longer honours the deprecated grpc.WithBlock dial option. If your code relied on Dial blocking until connected, follow the grpc-go anti-patterns guidance referenced in the changelog. Container images are published only as a multi-arch manifest; the per-architecture tags are gone, so a pinned v3.7.x-arm64 style tag will fail to pull.
Kubernetes Control Plane Impact: Watch Cache, Lists, and RangeStream
For a Kubernetes control plane, etcd 3.7 matters most through one interaction: how the kube-apiserver fills and refills its watch cache. The API server does not forward every client read to etcd. It keeps a per-resource watch cache, initialised by a full list from etcd and kept current by a watch, and serves most reads and watches from memory.
That cache is built from large list calls, and those calls are the expensive part. The Kubernetes v1.37 post identifies two moments when the API server reads big collections: cache initialisation at startup and cache re-initialisation after a failure such as a compacted revision. Resources with many objects, or with large objects such as Pods, make these reads spike memory on both sides.

Figure 2: The MVCC read path. The treeIndex maps keys to revisions in memory. Keys-only reads can stop there, while value reads go to bbolt. Compaction drops old revisions but only defragmentation shrinks the file.
How the API server will use RangeStream
Per the KEP, RangeStream is consumed in kube-apiserver behind a feature gate named EtcdRangeStream, which graduates to Beta in Kubernetes v1.37. The announcement for etcd 3.7 says it becomes available in v1.37. Both the Kubernetes and etcd sides must be new: the API server needs v1.37 with the gate enabled, and etcd must be 3.7 or later, because older members do not implement the RPC.
Beta in Kubernetes terms usually means the gate is on by default, but I did not confirm the default state in the sources I retrieved, so check the v1.37 feature-gate table for your distribution before assuming either way. Managed offerings and distributions such as K3s publish their own 1.37 notes and may differ.

Figure 3: A RangeStream exchange. The server pins a revision, streams adaptively sized chunks, and the client merges them. A compaction that removes the pinned revision aborts the stream and forces a retry.
The sequence in Figure 3 makes the failure mode visible. The stream is consistent because it reads one revision, and it is bounded because chunks are small. But a long stream against a busy cluster is a race against compaction. The KEP treats this as equivalent to the existing paginated-list compaction race: the watch cache treats ErrCompacted as an initialisation failure and retries.
Think about what that means at scale. If your compaction interval is short and a list takes minutes, the retry loop can repeat. Streaming lowers peak memory but does not by itself make a huge list fast. The sensible operating posture is to keep compaction at a cadence that comfortably exceeds your slowest observed list, and to watch the apiserver’s list latency alongside etcd’s etcd_mvcc_db_total_size_in_use_in_bytes.
What to expect on CPU and memory
The honest expectation is directional rather than numeric. The memory benefit applies to peaks during large reads, not to steady-state RSS dominated by the treeIndex and bbolt’s memory-mapped pages. The CPU claim in the announcement is about etcd members compared with 3.6, with no workload description published.
My working hypothesis, labelled as opinion, is that clusters with a few resource types holding most of the bytes, such as Pods in a big node pool or Custom Resources from operators, will see the clearest gain. Small clusters whose whole keyspace is a few megabytes will not notice. If you want a number you can trust, replay the same restart of kube-apiserver against a staging etcd on 3.6 and on 3.7 and compare peak RSS and time to cache sync. That experiment is cheap, and it is the only benchmark that reflects your object mix.
Compatibility and where Kubernetes pins etcd
Kubernetes releases are tested against a specific etcd minor. For production use, follow your distribution’s supported etcd version rather than jumping ahead, since kubeadm, managed services, and K3s each bundle and qualify their own. Running the API server and etcd out of lockstep is not forbidden, but you lose the testing that the Kubernetes release actually received. The RangeStream pairing also means that adopting 3.7 early gives you the cheaper keys-only and lease paths immediately, while the streaming benefit arrives only with the matching API server.
Operating etcd: fsync, Quotas, Compaction, Defragmentation
Disk latency, not CPU, is the dominant risk factor for etcd. Every committed write requires the write-ahead log to be flushed to stable storage on a quorum of members, so fsync latency sets the floor for write latency, and slow fsyncs delay heartbeats and trigger elections. The hardware guidance calls for at least 50 sequential IOPS for light use and 500 sequential IOPS (a typical local SSD or a high-performance virtualised block device) for heavy loads.
The same page gives sizing examples. A small cluster is under 100 clients, under 200 requests per second, and under 100 MB of data, roughly 50 Kubernetes nodes. A large cluster serves up to 1,500 clients and 10,000 requests per second with under 1 GB of data, roughly 1,000 nodes, and an extra-large one exceeds those, around 3,000 nodes. Recommended CPU is two to four cores for typical clusters and eight to sixteen dedicated cores for heavy ones, with 8 GB of memory typical and 16 to 64 GB for thousands of watchers and millions of keys.
Measure fsync, do not assume it
The metric to alert on is etcd_disk_wal_fsync_duration_seconds. Its 99th percentile should stay in single-digit milliseconds on healthy SSD-backed storage; the commonly cited guardrail from the etcd documentation is a 99th percentile under 10 ms. I am citing that threshold as widely published etcd guidance rather than a figure I re-verified on a specific page this run. Pair it with etcd_disk_backend_commit_duration_seconds, which tracks bbolt commits, and with etcd_server_leader_changes_seen_total, which turns “flaky” into a counted event.
Before production, run a disk test that resembles the WAL’s access pattern: small sequential writes, each followed by fdatasync. The fio profile used by the etcd community does exactly this.
# Synthetic WAL-like load: 22 MiB total, 2300-byte blocks, fdatasync after each write
fio --rw=write --ioengine=sync --fdatasync=1 \
--directory=/var/lib/etcd-test --size=22m --bs=2300 \
--name=etcd-wal-check
# Read the fsync/fdatasync latency percentiles in the output;
# the 99th percentile should be well under 10 ms.
On Linux, the tuning guide also suggests raising etcd’s I/O scheduling class with ionice -c2 -n0 -p $(pgrep etcd), so a noisy neighbour on the same device is less likely to starve the WAL. Better still is a dedicated device for the data directory. In cloud environments, provisioned-IOPS volumes beat burstable ones precisely because burst credits run out under sustained control-plane churn.
Space quota, compaction, and defragmentation
etcd protects itself with a backend space quota. Per the FAQ, the default is 2 GB and the suggested maximum is 8 GB; etcd warns at startup if you configure more. When the database exceeds the quota, etcd raises a NOSPACE alarm and the cluster goes read-and-delete only until you recover. The documented recovery order is: compact, defragment, then disarm the alarm.
# 1. Find the current revision and compact to it
rev=$(etcdctl endpoint status --write-out=json | \
jq -r '.[0].Status.header.revision')
etcdctl compact "$rev"
# 2. Defragment members one at a time (blocks that member while it runs)
etcdctl --endpoints=https://10.0.0.11:2379 defrag
etcdctl --endpoints=https://10.0.0.12:2379 defrag
etcdctl --endpoints=https://10.0.0.13:2379 defrag
# 3. Clear the alarm and verify
etcdctl alarm disarm
etcdctl endpoint status --cluster -w table
Two facts make this routine safer. Live defragmentation blocks reads and writes on the member while it rebuilds the file, so you defragment one member at a time, leader last, and never defrag --cluster on a busy production control plane unless you accept simultaneous short stalls. And compaction alone does not shrink the file; it frees pages inside bbolt for reuse. The gap between etcd_debugging_mvcc_db_total_size_in_bytes and etcd_mvcc_db_total_size_in_use_in_bytes is your reclaimable fragmentation.
Who compacts? In a Kubernetes cluster, the kube-apiserver issues compaction requests on a timer, so you normally do not configure etcd auto-compaction there. I believe the default interval is five minutes via --etcd-compaction-interval, but I could not retrieve that flag’s documentation this run, so confirm with kube-apiserver --help. For standalone etcd, use --auto-compaction-mode=periodic with --auto-compaction-retention=1h or the revision mode with a fixed revision count, which the maintenance guide describes.
Quorum sizing and topology
Use three members for most clusters and five where you must survive two simultaneous failures, for example across zones during a rolling maintenance window. Seven or more rarely helps because every write must reach more machines, and the FAQ states plainly that larger clusters sacrifice write performance. Never put members on the same failure domain: three members on one hypervisor host behave as one member with extra overhead.
Latency across zones is workable but not free. With members in three availability zones at, say, 1 to 2 ms round trip, the default 100 ms heartbeat and 1000 ms election timeout are generous. Stretching a cluster across regions at tens of milliseconds pushes you into the tuning guide’s territory, where you raise heartbeat and election timeouts together, and accept slower failover. For a multi-region Kubernetes estate, prefer separate clusters per region over one stretched control plane.
The etcd Upgrade Guide: From 3.6 to 3.7 Without Surprises
An etcd upgrade from 3.6 to 3.7 is a rolling upgrade, one member at a time, and it is reversible only until the last member has been replaced. The official upgrade guide states one hard prerequisite: the cluster must already be on v3.6.11 or later, and earlier 3.6 patch versions are not supported for rolling upgrade to 3.7.

Figure 4: The etcd upgrade path. Binary rollback is possible while the cluster is in a mixed-version state; once every member runs 3.7, recovery means a supported downgrade or a snapshot restore.
Pre-flight checks
Start with the facts the guide makes non-negotiable. Confirm the version on every member, confirm health, and remove anything that 3.7 deleted. Three categories matter: v2 usage, --experimental-* flags, and Go client code.
Experimental flags are gone entirely. Per the guide, you must replace each with its 3.6 equivalent or a feature-gate entry before upgrading, otherwise the member refuses to start with an unknown-flag error. Grep your static pod manifests, systemd units, and Helm values for the prefix. Check the 3.6 documentation for the current spelling of each flag you use rather than guessing, because the replacement is not always a one-to-one rename.
# Version and health on all members
etcdctl --endpoints=$EPS endpoint status --cluster -w table
etcdctl --endpoints=$EPS endpoint health --cluster
# Hunt for leftover experimental flags in static pod manifests
grep -R -- "--experimental-" /etc/kubernetes/manifests/ /etc/systemd/system/ 2>/dev/null
# Check for v2 data (v3.5.32 and 3.6.13 or newer ship the improved check)
etcdutl check v2store --data-dir /var/lib/etcd
Next, check the v2 situation. If the v2 store holds real data from an ancient deployment, follow the project’s v2 migration guidance first, because 3.7 cannot read it. Kubernetes clusters almost never have v2 data, but a surprising number have monitoring or discovery scripts that call v2 endpoints.
Backup before anything else
Take a snapshot from the leader, verify it, and store it off the member. The upgrade guide lists this as a step; I treat it as mandatory because after the last member upgrades it is your guaranteed way back.
ETCDCTL_API=3 etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
snapshot save /backup/etcd-pre-3.7-$(date +%F).db
etcdutl snapshot status /backup/etcd-pre-3.7-$(date +%F).db -w table
The status output reports hash, revision, total keys, and size. Record them. A snapshot you have never restored is a hypothesis, so schedule a restore into a scratch environment before upgrade day, and time it. Disk throughput guidance in the hardware page implies that about 10 MB/s recovers 100 MB in around 15 seconds, so a multi-gigabyte database restoring slowly is a sign of a storage problem.
The rolling procedure
The guide’s procedure is simple: stop one member, replace the binary, restart it with identical configuration, verify, and repeat. The cluster remains available with the reduced member count, so with three members you run at quorum-minimum during each step. That is exactly why you must check health between steps and never begin the next member until the previous one is healthy and caught up.
My own ordering recommendation, labelled as practice rather than documentation, is followers first and leader last. It avoids a leader election during the middle of the procedure. Find the leader with etcdctl endpoint status -w table and read the IS LEADER column. When you finally restart the leader, an election occurs, which is brief on a healthy cluster, but it lands at the end rather than at the start.
For kubeadm-style static pods, the upgrade is an image-tag change in the manifest, and kubelet restarts the pod. Use the multi-arch manifest tag, not an architecture-suffixed one, since 3.7 stopped publishing those. After each member, verify four things: the member reports the new version, endpoint health is green, the Raft index is converging across members, and the kube-apiserver shows no sustained error rate.
# After each member comes back
etcdctl --endpoints=$EPS endpoint status --cluster -w table
# VERSION should show 3.7.x for upgraded members, RAFT INDEX should be close
# together, and no member should show ERRORS.
Pick the current patch, v3.7.2, rather than v3.7.0. The 3.7.1 release fixed a case where a user granted read permission on one key could receive watch responses for every key starting from that key (advisory GHSA-xg4h-6gfc-h4m8), bounded an unrestricted io.ReadAll on the peer lease HTTP handler, and added timeouts for client HTTP headers and TLS handshakes. The 3.7.0 release itself fixed a CRL enforcement bypass on the gRPC listener when --listen-client-http-urls was set. The 3.7.2 release fixes an fsync omission on the snapshot directory when a member receives a snapshot, and bumps gRPC and OpenTelemetry for CVEs.
Rollback and downgrade
While at least one member still runs 3.6, you can roll the 3.7 binary back to 3.6 without a formal downgrade. After all members are on 3.7, binary rollback is not possible. Your options then are the official downgrade procedure, introduced in 3.6, or restoring the pre-upgrade snapshot into a fresh cluster, which means losing writes since the snapshot.
This asymmetry should shape your change window. Do not upgrade the last member until the first two have run under real load for a meaningful interval, and do not combine the etcd upgrade with a Kubernetes control-plane upgrade in the same hour. If something breaks, you want to know which change did it.
Restore drill, with the Kubernetes caveat
Restoring is where teams discover what they never tested. The restore command is etcdutl snapshot restore (the restore function lives in etcdutl, the offline utility, rather than etcdctl), and it creates a new data directory with a new cluster identity: the recovery guide states that the member ID and cluster ID are overwritten, so a restored member cannot accidentally rejoin the old cluster.
etcdutl snapshot restore /backup/etcd-pre-3.7-2026-10-06.db \
--name etcd-a \
--data-dir /var/lib/etcd-restored \
--initial-cluster etcd-a=https://10.0.0.11:2380,etcd-b=https://10.0.0.12:2380,etcd-c=https://10.0.0.13:2380 \
--initial-cluster-token etcd-restore-2026-10 \
--initial-advertise-peer-urls https://10.0.0.11:2380 \
--bump-revision 1000000 --mark-compacted
The two flags at the end are the Kubernetes-specific part. The recovery guide says that when restoring where controllers use watches, revision bumps are highly recommended: --bump-revision adds to the snapshot’s revision so revisions never go backwards relative to what clients have seen, and --mark-compacted marks all revisions compacted, which terminates watches and invalidates informer caches so clients relist. Without them, a controller may hold a resourceVersion newer than the restored store and behave incorrectly. The value 1000000 above is illustrative; choose a bump larger than the number of revisions your cluster could have produced since the snapshot.
Run the same restore command on every member with its own --name and peer URL, then start them together. Practise this on a copy, because the exact flags depend on your topology.
Trade-offs, Gotchas, and What Goes Wrong
The failures that take down control planes are repetitive, and none of them is new in 3.7. Knowing the pattern is more useful than memorising the release notes.
Slow disk masquerading as network trouble. A shared or burstable volume produces long fsync tails. The leader fails to send heartbeats on time, followers time out, an election starts, and the API server logs leader changes. The fix is storage, not timeouts. Raising the election timeout hides the symptom while lengthening failover.
Database full. A runaway controller that writes Events or ConfigMaps in a loop fills the quota, the NOSPACE alarm fires, and the cluster becomes read-and-delete only. Compaction and defragmentation are the recovery, but finding the writer is the repair. Count keys by prefix with a keys-only range, which in 3.7 is cheaper because it never reads values.
Defrag stall. Running defrag --cluster serially hides the fact that each member is unavailable while it rebuilds. On a three-member cluster that is survivable one at a time, and painful if two overlap with a failure. Never defragment during an upgrade.
Quorum arithmetic errors. Four members, or two zones holding three members as two plus one, are classic traps. With three members placed as two in one zone and one in another, losing the zone that holds two members leaves a single survivor, which is below the quorum of two, and the cluster stops. Place members across at least three failure domains so that no single domain holds a majority.
Large objects and large lists. etcd’s default maximum request size is about 1.5 MiB, a figure I am quoting from general etcd documentation rather than re-verifying this run. Objects near that size, and lists that return hundreds of megabytes, strain everything. RangeStream relieves the list memory spike, not the underlying bloat.
Embedded Go consumers. Operators and tools that import etcd client packages and ship their own binary can fail to compile or change behaviour after the protobuf migration and the non-blocking dial change. Pin client versions and test in CI before bumping.
Source disagreements. Secondary write-ups of 3.7 contain conflicting details. For example, one announcement page lists bbolt v1.5.1 and Go 1.26.4, while the changelog entry for 3.7.0 lists bbolt v1.5.0 and Go 1.26.5. When in doubt, trust the tagged changelog and verify with etcd --version, which prints the Go version the binary was built with.
There is also an honest limit on what this post can tell you. RangeStream’s Kubernetes performance numbers, the default state of FastLeaseKeepAlive, and bbolt’s new file-size controls are not documented with figures or defaults in the sources I retrieved. Those are the areas where you should benchmark on your own cluster.
Practical Recommendations
If you operate Kubernetes on etcd today, the order of operations is boring on purpose. First, get to 3.6.11 or later if you are not there, since that is the gate for the rolling upgrade, and retire any 3.4 members immediately because that branch no longer receives patches.
Second, upgrade staging to 3.7.2 and record two numbers on your own workload: peak kube-apiserver memory during a cold start, and time to watch-cache sync. Re-run them on 3.7 etcd with the API server unchanged, then again when your Kubernetes version brings RangeStream. Those measurements, not the announcement, justify the production change.
Third, put the control-plane hygiene in place before the upgrade rather than after: dedicated fast disks, alerts on fsync latency, leader changes and database size, off-host snapshots with a tested restore, and a runbook that distinguishes compaction from defragmentation.
- Run an odd member count, spread across three failure domains, with a dedicated SSD or NVMe data volume per member.
- Alert on 99th percentile WAL fsync latency, backend commit latency, leader-change rate, and database in-use size against quota.
- Keep the quota at the default 2 GB unless you have measured need, and never above 8 GB without a reason.
- Remove
--experimental-*flags and any v2 dependency before the upgrade, usingetcdutl check v2store. - Snapshot from the leader, verify with
etcdutl snapshot status, and rehearse restore with--bump-revisionand--mark-compacted. - Upgrade one member at a time, followers first, checking health between steps, and target v3.7.2.
- Defragment one member at a time, outside upgrade windows.
- Pin the multi-arch image tag and test any Go code that embeds etcd client packages.
Frequently Asked Questions
What is new in etcd 3.7?
The main additions are the RangeStream API for streaming large reads, key-only range requests served from the in-memory index, prioritised lease revocation, FastLeaseKeepAlive, Unix socket endpoints, and bootstrap entirely from the v3 store. It removes the v2 store and API, all --experimental-* flags, and v2discovery. It also moves to google.golang.org/protobuf and Go 1.26, and publishes only multi-arch container images.
Can I upgrade directly from etcd 3.5 to 3.7?
The official guide documents a rolling upgrade from 3.6, and requires the cluster to be on v3.6.11 or later first. Etcd has traditionally supported moving one minor version at a time, so go 3.5 to 3.6 and then 3.6 to 3.7, validating health between hops. I did not find a documented supported path that skips 3.6, so plan for two upgrades.
Is it safe to roll back after upgrading to etcd 3.7?
Only partly. While any member still runs 3.6, you can swap that member’s binary back without a formal downgrade. Once every member runs 3.7, binary rollback is not possible, and you must use the official downgrade procedure introduced in 3.6 or restore a pre-upgrade snapshot into a new cluster, losing writes made since the snapshot. Take and test a snapshot first.
How many etcd members does a Kubernetes cluster need?
Three for most clusters, five if you must tolerate two simultaneous failures. Quorum is floor(N/2)+1, so three members tolerate one failure and five tolerate two, while four members tolerate only one, the same as three. Larger clusters slow writes because each commit needs acknowledgement from more members. Spread members across independent failure domains, not just across hosts.
How do I back up and restore etcd?
Use etcdctl snapshot save against the leader with the certificate flags, verify with etcdutl snapshot status, and store the file off the member. Restore with etcdutl snapshot restore, which creates a new data directory and new cluster identity. For Kubernetes, add --bump-revision and --mark-compacted so watchers and informer caches relist rather than trusting stale revisions.
What is the difference between compaction and defragmentation in etcd?
Compaction discards old MVCC revisions below a chosen revision, freeing pages inside the bbolt file but not shrinking it. Defragmentation rewrites the file to return that free space to the filesystem, and blocks reads and writes on the member while it runs. If you hit the NOSPACE alarm, compact first, then defragment each member, then disarm the alarm.
Further Reading
- Cilium vs Calico for industrial Kubernetes networking: an ADR, on how the network fabric affects control-plane traffic.
- Cilium vs Calico, second ADR perspective for industrial clusters, a companion decision record.
- CNI comparison: Calico, Cilium, Flannel and Multus on Kubernetes, for choosing the data plane next to your control plane.
- Azure IoT Akri and the Kubernetes resource interface, on edge workloads that add resources to etcd.
- Announcing etcd v3.7.0 (etcd project) and the etcd 3.7 changelog.
- Kubernetes v1.37: etcd RangeStream cuts memory use on large list reads.
By Riju — about
