Kubernetes 1.37: Stable Metrics API and Rootless Kubelet in Beta

Kubernetes 1.37: Stable Metrics API and Rootless Kubelet in Beta

Kubernetes 1.37: Stable Metrics API and Rootless Kubelet in Beta

Two headline items in Kubernetes 1.37 sound like big operational events and are not. The Kubernetes 1.37 Metrics API is now stable, yet the schema did not change by a single field. The rootless kubelet is now beta, yet the feature gate behind it does almost nothing on its own. Both are real milestones, and both are easy to misread as “flip it on and win”.

The release, codenamed Garhwal, shipped on 26 August 2026 with 67 enhancements. Its value for platform teams is a contract on one side and a node-hardening path on the other. The contract lets autoscaling stand on a promise instead of a habit. The path lets you shrink what a container breakout can reach on the host.

This post is written as an architecture decision record. You will leave knowing exactly what graduated, what did not, which prerequisites bite, how to migrate safely, and when to say no.

What this covers: the verified facts of each graduation, the HPA and metrics-server gap, the rootless threat-model delta, migration steps, failure modes, and a decision matrix.

Context and Background

The Kubernetes release blog counts 67 enhancements in v1.37: 16 graduated to Stable, 23 to Beta, 27 entered Alpha, and one is a deprecation or removal. The project has already covered the scheduling and device stories on this site, so this post skips them. If you want those, read our 1.37 DRA and GPU device plugin analysis, the gang scheduling comparison against Volcano and Kueue, and the in-place pod resize rightsizing guide.

The two items here sit at opposite ends of the stack. The resource Metrics API, metrics.k8s.io, is a read-only aggregated API that reports CPU and memory for nodes and Pods. It powers kubectl top and resource-based autoscaling. It first appeared as alpha in v1.6 and reached beta in v1.8, then stayed in beta for nearly nine years, according to the release notes.

The rootless kubelet lives in the opposite place: the node. KEP-2033, authored by Akihiro Suda, describes running the kubelet, the container runtime, CNI plugins and kube-proxy as an unprivileged user inside a Linux user namespace. It was merged as alpha in v1.22 in 2021 after starting as an experiment in 2018. In v1.37 the KubeletInUserNamespace gate reached beta.

Why do these two belong in one post? Both are graduations of long-running, quietly load-bearing work, and both change your risk posture more than your feature list. The Kubernetes project has a stated goal of avoiding permanent beta APIs, and KEP-5207 cites KEP-1635 as requiring APIs to graduate or be deprecated rather than linger. Rootless mode, meanwhile, answers a steady drumbeat of container-escape CVEs that end in host root.

There is also a subtle link. A rootless node changes where the kubelet listens and how cgroups are delegated, and those are exactly the things the metrics pipeline depends on. We will come back to that in the walk-through.

One caution on sources. Several third-party summaries of 1.37 blur the two graduations together or claim more than the primary sources support, for example that HPA has already moved to the v1 API. Every fact below comes from the Kubernetes blog posts, the KEP text, or the Kubernetes documentation, and unverified points are flagged. Primary sources are linked in Further Reading.

What the Kubernetes 1.37 Metrics API Stable Graduation Actually Changes

Direct answer: Kubernetes 1.37 promotes metrics.k8s.io from v1beta1 to v1. The v1 API has the same resource types and fields as v1beta1, so this is a version graduation, not a data change. No feature gate is involved. Your metrics provider must serve v1.metrics.k8s.io, and the HPA controller still reads only v1beta1 in this release.

Kubernetes 1.37 Metrics API stable pipeline from cgroups through metrics-server to kubectl top and HPA

Figure 1: The resource metrics pipeline in 1.37. Both API versions can be registered side by side; the consumers differ in which one they read.

The diagram shows the path a number takes from a cgroup to a scaling decision. The kubelet embeds cAdvisor, which reads container statistics from cgroups. The kubelet exposes node-level resource metrics on its /metrics/resource endpoint. Metrics-server scrapes every kubelet (its README states a 15-second collection interval), keeps the result in memory, and serves it through the API aggregation layer. Clients then talk to the kube-apiserver, which proxies to the registered APIService.

What “stable” buys you

The blog states the meaning plainly: the API now carries the stability guarantees of a Kubernetes stable API. In practice that means the object shapes, NodeMetrics and PodMetrics, cannot be removed or incompatibly changed without going through the deprecation policy. The PodMetrics object still carries a per-container breakdown in its containers field.

The value is not new capability. Every controller, dashboard, and agent that reads this API for autoscaling or capacity signals now depends on a contract rather than on the project’s goodwill. For a platform team that builds tenant-facing autoscaling, that matters when you write your own compatibility policy and support matrix.

It is worth being blunt about the negative space. The API remains intentionally small. The blog says it is not a replacement for a full monitoring pipeline or for the custom metrics API, custom.metrics.k8s.io. Nothing in 1.37 adds a new signal such as network throughput, queue depth, or latency.

The two-version reality

Graduation works by adding an identical v1 alongside v1beta1. KEP-5207 lists as goals migrating the in-tree consumers, the HPA controller and kubectl top, to the stable version while keeping backward compatibility. It lists as a non-goal removing v1beta1 immediately: it keeps serving until formally deprecated. The release notes say the same: v1beta1 remains usable throughout the transition, in line with the API deprecation policy.

Because the API is served by an aggregated extension server, there is no in-tree implementation that “turns on” v1. The blog is explicit that you need no feature gate. For v1 to exist in your cluster, your chosen implementation must serve v1.metrics.k8s.io and you must register an associated APIService. During the transition, implementations should serve both versions so older clients keep working.

Two commands from the blog verify the state. kubectl get --raw /apis/metrics.k8s.io/ shows which versions the cluster serves, and kubectl get apiservice v1.metrics.k8s.io shows whether the v1 APIService is available. If the second returns not found, your provider does not yet serve v1, whatever your control plane version says.

Who reads which version in 1.37

This is the detail most summaries drop. The blog says kubectl top supports both versions: it prefers v1 when available and falls back to v1beta1 on clusters that do not serve v1. It then says the HorizontalPodAutoscaler controller currently supports only v1beta1. Discovery-based selection between the two is planned but is not available in Kubernetes v1.37.

That has a direct consequence. Even after your metrics-server serves v1, your HPAs keep reading v1beta1. So v1beta1 must stay registered and healthy in 1.37. Anyone who reads “graduated to stable” as “we can retire the beta endpoint” will break autoscaling.

The HPA gap also means the stable graduation is, for autoscaling, a forward-looking commitment. The controller migration is the follow-up, and the project has not published a release for it. Treat any claim that “HPA now uses v1” as unverified.

Rootless Kubelet in Beta: What KubeletInUserNamespace Really Does

The rootless story has a naming trap. Kubernetes has two different user namespace features, and the release blog goes out of its way to separate them. User namespaces for pods, set with hostUsers: false under the UserNamespacesSupport gate, went GA in v1.36. That puts pods in user namespaces while the node components still run as root.

KubeletInUserNamespace is the opposite. It is for running the node components themselves, meaning the kubelet, the CRI and OCI runtimes, CNI plugins and kube-proxy, as a non-root host user inside a user namespace. The two features do not conflict, and they can be combined to nest Kubernetes inside Kubernetes without resorting to a fully privileged: true pod. Our rootless containers comparison of Podman and Docker covers the container-runtime side of the same idea.

Rootless kubelet architecture showing user namespace, delegated cgroup v2, and port forwarder

Figure 3: A rootless node. Everything inside the box is owned by an unprivileged UID; the host contributes only cgroup delegation, a subordinate UID range, and a port forwarder.

The gate is boring on purpose

The blog calls the gate itself “quite boring”. It does not create a user namespace. The namespace has to be created outside Kubernetes, for example by rootless Docker, RootlessKit, unshare(1) or a runtime such as Podman. What the gate does is let the kubelet ignore permission errors that would otherwise abort startup.

Per the documentation, the kubelet ignores errors when setting these sysctl values: vm.overcommit_memory, vm.panic_on_oom, kernel.panic, kernel.panic_on_oops, kernel.keys.root_maxkeys and kernel.keys.root_maxbytes. Inside a user namespace it also ignores errors from opening /dev/kmsg, and it lets kube-proxy ignore an error when setting RLIMIT_NOFILE. That is essentially the whole code change.

The heavy lifting is elsewhere: the namespace, the delegated cgroup tree, the network plumbing, the runtime configuration. This matters for your planning. Enabling a gate is trivial; building a node image that boots cleanly into a user namespace is the project.

What changed from alpha to beta

The blog lists concrete changes. The gate is now enabled by default, but enabling it does not move a kubelet into a user namespace, so existing rootful clusters see no difference. Node objects now report whether they run in a user namespace through the runningInUserNamespace property, visible with kubectl get nodes -o yaml. The KEP describes this as a new RunningInUserNamespace field on NodeSystemInfo.

Administrators can use that property to label or taint rootless nodes and keep workloads that need real root, such as some CNI plugin installers, away from them. Kubernetes’ own CI also now runs node conformance tests on a rootless cluster (ci-kubernetes-e2e-kind-rootless). Three related improvements landed outside the gate: Linux kernel 6.3 added idmapped tmpfs support, Kubernetes 1.33 enabled UserNamespacesSupport by default, and containerd 2.1 added writable cgroups.

The prerequisites, verified

The documentation lists them under “Before you begin”. Your server must be at v1.22 or later. You need cgroup v2, systemd with a user session, several sysctl values configured for your distribution, and your unprivileged user listed in /etc/subuid and /etc/subgid. Then you enable the feature gate.

cgroup v1 is not supported. That aligns with a separate 1.37 note: since v1.35 the kubelet’s failCgroupV1 setting defaults to true, so the kubelet fails to start on cgroup v1 nodes unless overridden, and cgroup v1 removal is planned for a future release (KEP-5573). Rootless nodes have no v1 escape hatch at all.

For the runtime, the documentation says containerd supports running its CRI plugin in a user namespace since 1.4, and CRI-O since 1.22, the latter requiring _CRIO_ROOTLESS=1. Our containerd versus CRI-O comparison is a useful companion for choosing between them.

The documented kubelet configuration is short:

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
  KubeletInUserNamespace: true
# cgroupfs delegated by systemd, not the systemd driver
cgroupDriver: "cgroupfs"

The documented containerd settings for a user namespace disable AppArmor, set restrict_oom_score_adj = true, set disable_hugetlb_controller = true because systemd does not delegate the hugetlb controller, use the fuse-overlayfs snapshotter (non-FUSE overlayfs is possible on kernel 5.11 or later but needs SELinux disabled), and set SystemdCgroup = false.

kube-proxy needs its own configuration: mode: iptables (or userspace), with conntrack.maxPerCore: 0, tcpEstablishedTimeout: 0s, and tcpCloseWaitTimeout: 0s to skip conntrack sysctls it cannot write.

The caveats list is the real product spec

Do not adopt on the strength of the headline. The KEP and docs list constraints that decide whether a node pool can go rootless. Non-local volume drivers such as NFS and iSCSI do not work, because a user namespace supports only tmpfs, bind mounts and FUSE filesystems. Local volumes, emptyDir, hostPath, configMap, secret and downwardAPI are known to work.

Some CNI plugins do not work; Flannel with VXLAN is the known-good choice. The KEP adds that hugepages cannot be supported, because systemd does not delegate the hugetlb controller. AppArmor is unsupported, so a Pod requesting an AppArmor profile fails. Some pod-level sysctls fail with EPERM. The runAsUser range is limited by /etc/subuid. NodePorts below 1024 cannot be exposed by default, which is a non-issue with the 30000 to 32767 default range. The KEP also notes that network performance is limited by slirp4netns overhead unless you use the lxc-user-nic setuid helper, a trade against the “no privileged binaries” goal.

Finally, the kubelet’s port 10250 and all NodePorts live inside the node’s network namespace, so an external port forwarder such as RootlessKit or socat must expose them. That detail returns in the next section, because metrics-server needs to reach 10250.

Threat Model: What Rootless Actually Removes

A threat model needs a clear adversary. Here the adversary is a workload that escapes its container, or an attacker who exploits a bug in a node component, and the question is what they own afterwards. On a conventional node the answer is host root. On a rootless node it is the unprivileged account that owns the user namespace.

The blog motivates the feature with a list of real vulnerabilities in node components that led to host root: CVE-2022-0811 (CRI-O “cr8escape”, arbitrary sysctls including kernel.core_pattern), CVE-2023-27561 (runc masked-path bypass), CVE-2024-10220 (kubelet command execution via gitRepo volumes), CVE-2025-31133 (runc bind-mounting attacker-controlled paths into procfs) and CVE-2026-53488 (containerd executing commands via crafted image labels). Rootless mode does not prevent these bugs. It caps the damage.

What you gain, and what you keep

The blog says that by running node components in a user namespace, potential damage is confined to the non-root user’s account. It also states that an attacker cannot conceal their intrusion by modifying the kernel, the boot loader or the firmware, since those need real root. That is a meaningful detection and recovery property: a rebuild-from-image workflow can trust the lower layers.

Just as important is what you keep exposed. The blog states that user namespaces are not effective at mitigating vulnerabilities in the kernel itself, and recommends pairing them with hardening such as seccomp. The KEP says so directly: a bug in the kernel’s user namespace implementation could let root in the namespace reach real root, and it prefers sandbox technologies like gVisor for pods where that risk matters. For the sandbox comparison, see our Firecracker, gVisor and Kata analysis.

The delta in table form

Scenario Rootful node Rootless node
runc or containerd escape to node component Host root Unprivileged user’s account
Kubelet command-execution bug Host root Unprivileged user’s account
Attacker persistence in kernel, boot loader, firmware Possible with host root Not reachable without a kernel bug
Kernel user namespace vulnerability Host root Host root possible
Privileged pods Real root on host Fake root, confined to the namespace
AppArmor profile for pods Supported Unsupported

The last two rows show the trade. Privileged pods are far less dangerous, and that is a feature, but many node agents that legitimately need host privileges cannot run on such a node at all. The taint on runningInUserNamespace is how you route them elsewhere.

One more delta: the attack surface shifts to the pieces you must add. A port forwarder, a setuid lxc-user-nic helper if you choose it, newuidmap and newgidmap binaries, and the delegation config are all new trust points. The net is still favorable for most threat models, but you should count them, not ignore them.

Migration Walk-Through: From v1beta1 to v1 Without Breaking Autoscaling

The safe order is control plane first, provider second, consumers third, cleanup never (yet). Each step has a check that either passes or tells you to stop.

Sequence of kubectl top and HPA discovering metrics.k8s.io versions and reading from metrics-server

Figure 2: kubectl top negotiates a version and falls back; the HPA controller in 1.37 goes straight to v1beta1, which is why the beta APIService must stay registered.

Step 1: snapshot the baseline

Before touching anything, record what works. If an HPA already shows <unknown> targets, fix that first, because you cannot audit a change against a broken baseline.

kubectl get apiservice | grep metrics.k8s.io
kubectl get --raw /apis/metrics.k8s.io/ | jq .
kubectl top nodes
kubectl get hpa -A

Step 2: upgrade the control plane and confirm beta survives

After the 1.37 upgrade, v1beta1.metrics.k8s.io must still report Available: True. That is expected, since the beta API remains available in 1.37. Nothing changes for consumers at this point, which is the point of the graduation design. The API is read-only and holds no persisted state, so KEP-5207 describes upgrade and downgrade as symmetric with no data migration.

Step 3: upgrade the provider, but check that a release exists

Here I have to flag an open item. The blog says the provider must serve v1 and register the APIService. At the time of writing, the metrics-server compatibility matrix in its repository lists 0.9.x, released 13 July 2026, serving metrics.k8s.io/v1beta1 only. A pull request adding v1.metrics.k8s.io support (PR 1855) was described by its author as remaining a draft until 1.37 shipped. I could not confirm a tagged release that serves v1.

So the practical guidance is: read your provider’s release notes and check the actual APIService, rather than assuming. If you run metrics-server, look for a release that registers v1.metrics.k8s.io. If you run a managed cluster, your vendor may ship its own build. Until v1 is available, staying on v1beta1 is fully supported.

Step 4: probe both versions from the client side

kubectl get --raw /apis/metrics.k8s.io/v1beta1/nodes | head -c 300
kubectl get --raw /apis/metrics.k8s.io/v1/nodes | head -c 300
kubectl get apiservice v1.metrics.k8s.io

The first must always succeed. The second succeeds only where a v1-capable provider is registered. A 404 on the second while kubectl top works means you are on fallback, which is fine until your provider catches up.

Step 5: verify HPA behavior, not version strings

Since the HPA reads v1beta1 in 1.37, the check is behavioral. Pick a non-critical HPA-backed workload and confirm the loop closes.

kubectl describe hpa web -n prod | grep -iE "ScalingActive|FailedGetResourceMetric|metrics"

You want ScalingActive: True with no FailedGetResourceMetric events. Right after a metrics-server restart, targets can briefly read <unknown> while it collects its first window of samples; the docs example shows a 30-second window per sample. Anything persistent beyond a few minutes points at the APIService or kubelet connectivity.

Step 6: soak, and do not delete v1beta1

Leave both registered. No release yet reads v1 in the HPA, and older agents and pinned clients still prefer beta. The most common self-inflicted outage in an API graduation is premature cleanup by someone who reads “deprecated” as “dead”. Retire beta only after a documented controller and provider combination says it is safe, then confirm zero beta traffic through API server request metrics before removing it.

Sizing metrics-server: worked numbers

Graduation does not change resource needs, but if you are touching the deployment anyway, use the numbers the project publishes. Its README states default requests of 100m CPU and 200MiB memory, sized for clusters up to 100 nodes, 70 pods per node and 100 Deployments with HPAs. For larger clusters it says to allocate an additional 1m core and 2MiB per node. I read that as per node beyond 100, though the README is not explicit.

Cluster size CPU request Memory request Basis
100 nodes 100m 200MiB README default
500 nodes about 500m about 1,000MiB Illustrative: default plus 400 nodes at 1m and 2MiB
1,000 nodes about 1,000m about 1,800MiB Illustrative: default plus 900 nodes at 1m and 2MiB

These are derived estimates, not benchmarks. High availability needs at least two nodes for scheduling and, per the README, the kube-apiserver flag --enable-aggregator-routing=true to balance requests across replicas.

HPA, VPA, KEDA and the Wider Autoscaling Picture

The resource Metrics API feeds two families of consumers, and the graduation touches them differently.

HPA

The HPA controller reads v1beta1 in 1.37, so nothing changes in its behavior. The release also brings HPA scale to zero to beta, enabled by default. It works only for object or external metrics, with minReplicas: 0, and the release notes say scaling to zero on CPU and memory is not supported because those metrics depend on active Pods. That is a separate feature and a common source of confusion with the Metrics API news. When an HPA holds a workload at zero, it records a ScaledToZero condition to distinguish it from a manual zero.

VPA

The Kubernetes docs say the VerticalPodAutoscaler uses data from the Metrics API. The 1.37 blog says HPA and VPA both continue to function. I could not verify which API version the VPA recommender client currently requests, so treat VPA the same way as HPA: keep v1beta1 registered and verify after upgrade. Also note that in-place pod resize interacts with VPA modes; our in-place pod resize guide covers that.

KEDA and custom metrics

Event-driven scalers do not read the resource Metrics API for their event signals. They feed external metrics into the HPA. A stable metrics.k8s.io therefore does not remove the need for the custom or external metrics path if you scale on queue depth, request rate or latency. Our KEDA event-driven autoscaling architecture explains that path, and the KEDA 2.21 migration note covers a security-driven upgrade. For node scaling, see the Karpenter versus Cluster Autoscaler comparison.

The interpretation I would defend: use the stable Metrics API for what it is, a small contract for CPU and memory, and put every richer signal on a pipeline designed for it. Stability makes that split cleaner, not obsolete.

Where the Two Graduations Collide: Metrics on Rootless Nodes

If you adopt both, a specific interaction appears. Metrics-server needs to reach each kubelet’s port, 10250 by default, at the node address published in the Node object. Its README also lists requirements such as webhook authentication and authorization on nodes, and a container runtime that implements the container metrics RPCs or has cAdvisor support.

On a rootless node the kubelet’s port lives inside the node’s network namespace. The documentation says ports such as 10250 and NodePorts must be exposed to the host with an external port forwarder. If that forwarder is missing or misconfigured, the kubelet is healthy but unreachable, and metrics-server marks the node as failing. HPAs then lose their data. The symptom looks like a Metrics API problem when it is a node networking problem.

Second, cgroup accounting depends on the delegated subtree. The kubelet’s cAdvisor reads cgroups, and on a rootless node it can only see the subtree systemd delegated to the user. My inference, not a documented statement, is that a mis-delegated tree will produce missing or partial per-container stats. A pilot should include a kubectl top pods comparison against a rootful node running the same workload.

Third, hugepages. Because the hugetlb controller is not delegated, hugepage-backed workloads do not belong on rootless nodes, which is one more reason to keep them on a separate pool.

Memory QoS and cgroup v2

The 1.37 release moves memory QoS with cgroup v2 to beta. It uses memory controls such as memory.min to protect requested memory from reclamation. Rootless nodes are cgroup v2 by construction, but memory QoS depends on the kubelet being able to write those controls in the delegated subtree. Check this in your pilot rather than assuming, because the docs do not say either way.

A Rollout Plan for the Rootless Kubelet

The KEP is candid about the rollout model: rolling out means recreating a new node instance in a user namespace, and rolling back means recreating a node instance too. The gate itself can be disabled, but that does not move a running kubelet out of its namespace. This shapes everything else. Treat rootless as an immutable node-image change, not a config toggle.

Build the node image

Start from a distribution with cgroup v2 and systemd user sessions. Provision a dedicated unprivileged user, allocate its /etc/subuid and /etc/subgid range, and enable delegation of the cgroup controllers you need. The KEP names the two typical failures: subuids not allocated, and cgroup v2 delegation not enabled. Both show up as a node that never registers.

Then configure the runtime with the documented rootless settings, the kubelet with the gate and the cgroupfs driver, and kube-proxy with the conntrack overrides. Add a port forwarder for 10250 and NodePorts. Choose a CNI from the known-good set, which today means Flannel VXLAN, and validate anything else yourself.

Label and taint from the start

Use the runningInUserNamespace node property to drive a label or taint so only tolerant workloads land there. Node agents that need real root, such as some CNI installers and privileged DaemonSets, must not schedule on the pool. Our multi-tenancy architecture guide shows where a rootless pool fits as one isolation layer among several.

Pilot with realistic workloads

Pick stateless services that use emptyDir and secrets and not NFS. Run your admission policies, your observability agents and your CNI through the pool. Compare pod startup time and network throughput against a rootful pool, since slirp4netns-style paths can cost performance. No published benchmark is cited here, so measure yourself.

Decide per pool, not per cluster

The gate default being on means nothing changes at cluster level. The decision is per node pool. Multi-tenant pools running untrusted or semi-trusted code gain the most. Pools running storage-heavy, privileged, or hugepage workloads gain the least and lose the most.

Decision flow for adopting the Metrics API v1 and the rootless kubelet in Kubernetes 1.37

Figure 4: A decision flow for the 1.37 upgrade. The Metrics API branch is about verification; the rootless branch is about whether your node pool can tolerate the constraints.

Decision Record and Adoption Matrix

Framed as an ADR, the two decisions are independent. Record them separately so a delay on one does not block the other.

Decision 1: Metrics API. Upgrade to 1.37 on your normal schedule. Register both API versions once your provider supports v1, keep v1beta1 available, and do not change autoscaling configuration. Consequence: no immediate benefit, a lower long-term risk of API removal, and a small verification burden.

Decision 2: Rootless kubelet. Pilot on a dedicated node pool for multi-tenant or untrusted workloads, and leave the rest rootful. Consequence: a smaller blast radius for node-component escapes, offset by a rebuilt node image, a constrained CNI and storage set, and new components to maintain.

Situation Metrics API v1 Rootless kubelet Reasoning
Single-tenant cluster, standard workloads Verify only Skip for now No new capability; rootless cost exceeds benefit
Multi-tenant platform running customer code Verify, then prefer v1 in new tooling Pilot on a tainted pool Blast radius reduction is the point
Storage-heavy stateful pools with NFS or block volumes Verify only Do not adopt Unsupported volume types
GPU or hugepage pools Verify only Do not adopt Hugepages unsupported, device stacks need privileges
Edge or lab clusters on shared hosts Verify only Strong candidate No root needed on the host, easier local isolation
Cluster with heavy custom or external metrics scaling Unchanged Optional Resource API is not the constraint
CI or ephemeral test clusters Verify only Strong candidate kind, minikube and k3s already support it

The docs list several distributions as ways to try it: kind, minikube and k3s in rootless mode, and the third-party Usernetes project. For edge footprints, see our K3s edge production guide.

Trade-offs, Gotchas, and What Goes Wrong

These are the failure modes I would plan around, ordered roughly by how likely they are to page someone.

Retiring v1beta1 too early. The HPA controller reads only v1beta1 in 1.37. Removing that APIService, or letting a provider upgrade drop it, blinds every resource-based HPA. Symptoms are FailedGetResourceMetric events and <unknown> targets, and workloads stop scaling in either direction.

Assuming the control plane brings v1. The API server can offer the group version only if a provider registers it. Until then v1.metrics.k8s.io returns nothing. Clients that hardcode v1 will fail against clusters that have not upgraded their provider, so new tooling should discover versions and fall back the way kubectl top does.

Version skew in mixed fleets. During a rolling fleet upgrade, some clusters serve both versions and some only beta. Automation that assumes a uniform state breaks in the middle. Write checks that treat “beta only” as valid.

Rootless node that registers but is invisible. A missing port forwarder makes the kubelet unreachable. Logs, exec, port-forward and metrics all depend on it. This is the most likely bring-up failure after the two the KEP names, missing subuid ranges and missing cgroup delegation.

CNI and CSI surprises. The blog warns that caveats may break compatibility with specific CNI and CSI drivers. If your CNI installer needs real root, it fails or, worse, half-installs. Tainting rootless nodes and testing each DaemonSet is cheap insurance.

False confidence about the kernel. Rootless mode does not defend against kernel bugs. Teams that adopt it and relax seccomp, sandboxing or patch cadence lose more than they gain. The blog and KEP both say to combine it with traditional hardening.

Rollback is a rebuild. The gate can be turned off, but a kubelet in a user namespace stays there until the node is recreated. That makes rollback slow, so canary with a small pool and keep the rootful image available.

Beta is on by default, adoption is not. Some reports describe the beta as “enabled by default” as if it silently hardens nodes. It does not. The gate only relaxes error handling; a namespace must be provided by you.

Unpublished performance data. I found no official benchmark of rootless kubelet overhead. The KEP flags slirp4netns networking overhead qualitatively. Expect to measure your own numbers, and do not trust third-party percentages that lack a methodology.

Practical Recommendations

Treat the release as two separate work items, and time-box both. The Metrics API item is a half-day audit per cluster. The rootless item is a quarter-long pilot if you have a real multi-tenant need, and a “not now” if you do not.

For the Metrics API, the bar for done is that kubectl top works, every HPA reports ScalingActive: True, both versions are registered where your provider supports it, and no automation hardcodes v1. For rootless, the bar is a tainted pilot pool running real workloads for at least a full upgrade cycle, with the node image reproducible from code.

Checklist:

  • Snapshot APIService state, kubectl top output and HPA status before upgrading.
  • Upgrade the control plane, then confirm v1beta1.metrics.k8s.io is still available.
  • Check your metrics provider’s release notes for v1.metrics.k8s.io support, and register both versions when it exists.
  • Keep v1beta1 until a documented HPA and provider combination reads v1.
  • Make new tooling discover versions and prefer v1 with fallback.
  • Decide rootless per node pool, and never on pools with NFS, block volumes, hugepages or privileged node agents.
  • Bake subuid ranges, cgroup v2 delegation, the runtime config and a port forwarder into the node image.
  • Label or taint on runningInUserNamespace, and keep seccomp and kernel patching in force.
  • Plan cgroup v2 migration for any remaining v1 nodes, because v1 removal is on the roadmap.
  • Re-check the KEP for GA criteria before committing a production timeline.

On timing, the KEP’s GA criteria say to promote after at least two releases in beta, assuming no negative feedback. By that rule the earliest GA would be v1.39, which is my arithmetic and not a project commitment. The blog says only “a future release”.

Frequently Asked Questions

What is the Kubernetes 1.37 Metrics API stable change?

Kubernetes 1.37 promotes metrics.k8s.io to v1. The v1 API has the same resource types and fields as v1beta1, so it is a version graduation with no change to what is collected or returned. It gives the API stable-API guarantees. No feature gate exists; your metrics provider, such as metrics-server, must serve v1.metrics.k8s.io and register an APIService.

Do I need to change my HPA after upgrading to 1.37?

No. The HorizontalPodAutoscaler controller currently supports only v1beta1, and discovery-based selection between v1 and v1beta1 is planned but not available in 1.37. Your existing HPAs keep working, provided the v1beta1 APIService stays registered and available. After upgrading, confirm each HPA shows ScalingActive: True and no FailedGetResourceMetric events, rather than inspecting API version strings.

Is v1beta1 of the Metrics API being removed?

Not in 1.37. The release notes say v1beta1 remains usable throughout the transition, in line with the API deprecation policy, and KEP-5207 lists immediate removal as a non-goal. Providers should serve both versions for now. Any future removal would follow the standard deprecation process, so watch release notes, and do not delete the beta endpoint early because the HPA still reads it.

Is the rootless kubelet enabled by default in 1.37?

The KubeletInUserNamespace feature gate is on by default in beta, but enabling it does not put the kubelet in a user namespace, so existing clusters see no change. You still create the namespace yourself, for example with rootless Docker, RootlessKit or unshare. The gate mainly lets the kubelet ignore sysctl and /dev/kmsg permission errors that would otherwise stop it starting.

What are the prerequisites for running the kubelet rootless?

The documentation lists cgroup v2, systemd with a user session, suitable sysctl values for your distribution, and your unprivileged user in /etc/subuid and /etc/subgid, plus the feature gate. It also needs a runtime configured for a user namespace, containerd 1.4 or later or CRI-O 1.22 or later, and a port forwarder for the kubelet port. cgroup v1 is not supported.

How is this different from user namespaces for pods?

User namespaces for pods use hostUsers: false with the UserNamespacesSupport gate, GA since v1.36. They isolate pods while the kubelet and runtime still run as root. KubeletInUserNamespace runs the node components themselves as a non-root user. They do not conflict and can be combined, for example to nest Kubernetes inside Kubernetes without a fully privileged pod.

Further Reading

Internal:

External primary sources:

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *