Kubernetes 1.37 Memory QoS and Pod-Level Resource Managers Beta Explained

Kubernetes 1.37 Memory QoS and Pod-Level Resource Managers Beta Explained

Kubernetes 1.37 Memory QoS and Pod-Level Resource Managers Beta Explained

Most Kubernetes memory incidents share a shape. A container grows past what anyone planned, nothing happens for a while, and then the kernel’s out-of-memory (OOM) killer arrives with no warning and takes the process out at memory.max. The cluster had plenty of signals before that moment, but the kernel was never told which memory mattered. Kubernetes Memory QoS is the feature that tries to tell it, by writing the pod’s requests and limits into the cgroup v2 memory controller as protection and throttling thresholds rather than a single hard wall.

Kubernetes 1.37, released on 26 August 2026 under the name Garhwal, promotes Memory QoS to beta, and its sibling feature Pod-Level Resource Managers to beta too. The headline “now enabled by default” is easy to misread, though, and the safe reading is more interesting than the alarming one. This post explains exactly what the kubelet writes into each cgroup file, why the default behaviour was deliberately made inert, how throttling can quietly hurt latency, and what the pod-level managers change for NUMA-sensitive workloads.

What this covers: the cgroup v2 memory interfaces, the 1.37 beta defaults and the change that protects upgrades, the exact memory.high formula with worked numbers, tiered reservation and its node-wide limitation, throttling failure modes, pod-level resource managers, a rollout plan with an observability checklist, and an FAQ.

Context and Background

Before cgroup v2, Kubernetes had essentially one memory control: the limit. A container’s limits.memory became memory.limit_in_bytes under cgroup v1, and the request influenced only scheduling and, indirectly, eviction ranking. The kernel had no idea that a database container with a 4 GiB request deserved more protection from reclaim than a batch job with none. When the node ran short, the kernel reclaimed pages from whichever cgroups its generic heuristics preferred, and Kubernetes layered kubelet-driven eviction on top.

cgroup v2 changed what is possible. It provides a single unified hierarchy, pressure stall information (PSI), and four distinct memory knobs rather than one: memory.min, memory.low, memory.high and memory.max. The Kubernetes documentation on cgroup v2 lists the requirements: a distribution that enables cgroup v2, Linux kernel 5.8 or later, a runtime that supports it (containerd 1.4+ or CRI-O 1.20+), and the systemd cgroup driver. It also records that cgroup v1 has been deprecated since Kubernetes 1.35, with the kubelet refusing to start on a cgroup v1 node by default unless failCgroupV1 is set to false. Memory QoS only exists on the v2 side of that line.

Memory QoS itself is old. KEP-2570 introduced it as alpha in Kubernetes 1.22, and it sat there for years because the first design was too blunt. Enabling the gate wrote memory.min for every container with a memory request, which hard-locks that memory against reclaim. On a node whose Burstable pods requested most of RAM, that left no headroom for the kernel, system daemons or BestEffort workloads. Kubernetes 1.36 reworked the design with tiered reservation, and 1.37 finished the job by changing defaults so that graduation would not break running clusters.

That history is why the current feature looks the way it does. It is deliberately conservative, opt-in at the behaviour level even though the gate is on. If you operate edge fleets with small-memory nodes, such as the k3s clusters we cover in the edge Kubernetes production guide, the margin for error on memory protection is far smaller than on a 256 GiB cloud node, which makes understanding the exact semantics worth the effort.

The primary sources for everything below are the Kubernetes blog posts Memory QoS Graduates to Beta and Pod-Level Resource Managers graduated to Beta, the Kubernetes pod QoS documentation, and the Linux kernel cgroup v2 documentation.

How Kubernetes Memory QoS Maps Requests and Limits to cgroup v2

Kubernetes Memory QoS translates a pod’s memory requests and limits into cgroup v2 settings: memory.min gives Guaranteed pods hard protection, memory.low gives Burstable pods soft protection, memory.high throttles Burstable containers before they hit memory.max, the unchanged hard limit. In 1.37 the feature gate is on by default, but throttling and reservation stay off until you configure them.

Kubernetes Memory QoS mapping pod requests and limits to cgroup v2 memory interfaces

Figure 1: How the kubelet turns requests and limits into the four cgroup v2 memory files, by QoS class.

The diagram shows the flow in one pass. The kubelet reads each pod’s requests and limits, classifies the pod as Guaranteed, Burstable or BestEffort, and writes values into the pod, QoS-class and root kubepods cgroups. The kernel then enforces those values continuously, long before the kubelet’s own eviction logic would notice a problem. The important consequence is that protection and throttling happen inside the reclaim path, at page granularity, rather than as an after-the-fact kill.

The four knobs and what the kernel actually promises

The kernel documentation is precise about each file, and the differences matter more than the names suggest.

memory.min is hard protection. If a cgroup’s usage is within its effective min boundary, its memory will not be reclaimed under any conditions, and if no unprotected reclaimable memory remains, the OOM killer is invoked against someone else. The kernel documentation explicitly discourages protecting more memory than is generally available, warning of constant OOMs.

memory.low is best-effort protection. Memory within the boundary is not reclaimed unless there is no reclaimable memory left in unprotected cgroups. Above the boundary, pages are reclaimed proportionally to the overage. It is a preference, not a promise.

memory.high is a throttle, not a limit. When usage goes over it, the cgroup’s processes are throttled and put under heavy reclaim pressure. Crucially, the kernel documentation states that going over the high limit never invokes the OOM killer, and that under extreme conditions the limit may be breached. It is intended for scenarios where an external process watches the cgroup and relieves pressure.

memory.max is the hard limit. When usage reaches it and cannot be reduced, the OOM killer runs inside that cgroup. This is the only one of the four that Kubernetes has always set, from limits.memory.

Put together, the four files describe a gradient: protected floor, soft floor, throttle ceiling, kill ceiling. Classic Kubernetes only drew the last line.

What changed in Kubernetes 1.37

The beta announcement, written by Qi Wang and Sohan Kunkerkar of Red Hat, makes three things clear. First, the MemoryQoS feature gate is beta and on by default, so every 1.37 kubelet has it enabled with no configuration change. Second, enabling it is safe because the default kubelet configuration writes no memory.high, memory.min or memory.low values at all. Third, you opt into behaviours through two kubelet configuration fields.

The first field is memoryThrottlingFactor, a value between 0 and 1 that enables memory.high on Burstable and BestEffort containers. In 1.37 its default is null, meaning no throttling. In the alpha releases it defaulted to 0.9, so merely turning the gate on made the kubelet set memory.high everywhere. That was changed because, with the gate now on by default, an automatic memory.high could throttle workloads that previously ran unthrottled. Making the default null means upgrading to 1.37 does not change runtime behaviour.

The second field is memoryReservationPolicy, with values None (default) and TieredReservation. The latter writes memory.min and memory.low according to QoS class.

There is one upgrade subtlety worth underlining. If your kubelet configuration file already contains an explicit memoryThrottlingFactor, that value is preserved and throttling continues exactly as before. If it does not, you inherit the new null default and the kubelet stops setting memory.high. Anyone who relied on the 1.36 alpha behaviour, where the factor silently defaulted to 0.9 once the gate was on, will find throttling gone after the upgrade unless they add the field explicitly.

The mapping by QoS class

The Kubernetes 1.36 tiered-protection post publishes the clearest table of what is written when memoryReservationPolicy: TieredReservation is set. I reproduce it here in condensed form. The 1.37 beta post confirms the core of it: with TieredReservation, every Guaranteed pod gets memory.min and every Burstable pod gets memory.low.

QoS class memory.min memory.low memory.high memory.max
Guaranteed requests.memory (hard) not set not set, since requests equal limits limits.memory
Burstable not set requests.memory (soft) from formula with throttling factor limits.memory, if specified
BestEffort not set not set calculated from node allocatable memory not set

Two details deserve attention. A Guaranteed pod never gets memory.high because its request and limit are identical, so there is no gap to throttle within. And a BestEffort pod has no request to protect, so its memory stays fully reclaimable, which is exactly what you want when the kernel needs to free pages for someone more important.

The hierarchy also has a hard rule. cgroup v2 requires that a parent’s protection be at least as large as the sum of its children’s, so the kubelet sets memory.min on the root kubepods cgroup to the sum of Guaranteed and Burstable requests, and memory.low on the Burstable QoS cgroup to the sum of Burstable requests. The kubelet manages pod and QoS-level cgroups itself through the runc libcontainer library, while container-level cgroups belong to the runtime, which is why the choice between containerd and CRI-O still touches this feature’s correctness.

Reading the memory.high Formula and Tiered Reservation in Practice

The mapping above is conceptual. What operators need is arithmetic they can check against /sys/fs/cgroup on a real node. The Kubernetes pod QoS documentation gives the throttling formula for Burstable containers:

memory.high = requests + memoryThrottlingFactor * (limits - requests)

If a Burstable container has no memory limit, node allocatable memory stands in for the limit. The documentation’s own example is a container with a 256 MiB request and a 1 GiB limit, which gets memory.high of roughly 947 MiB at the 0.9 factor. You can verify that: 256 + 0.9 x (1024 – 256) = 947.2 MiB.

Worked examples with the factor (illustrative)

The following numbers are illustrative arithmetic, not measurements. They show how the throttle point moves as you change the factor on a container requesting 2 GiB with a 4 GiB limit.

memoryThrottlingFactor memory.high Headroom between memory.high and memory.max
0.5 3.0 GiB 1.0 GiB
0.8 3.6 GiB 0.4 GiB
0.9 3.8 GiB roughly 205 MiB
0.95 3.9 GiB roughly 102 MiB

The pattern is that a high factor gives the workload nearly all of its limit before throttling starts, but leaves a thin band in which the kernel can reclaim before memory.max triggers an OOM kill. A low factor throttles early and often. Neither is universally right. The factor is a lever between “reclaim pressure arrives late and sharply” and “reclaim pressure arrives early and constantly.”

The unbounded case is subtler. A Burstable container with a 1 GiB request and no limit, on a node with, say, 30 GiB allocatable (illustrative), gets a memory.high of 1 + 0.9 x 29 = 27.1 GiB under the documented rule. That is a very high throttle point, which means for no-limit containers the feature behaves almost like a guard against one pod consuming the entire node, rather than a tight governor. For BestEffort containers, the 1.36 post says only that memory.high is “calculated based on node allocatable memory”; I could not find the exact expression in the sources I fetched, so check the kubelet source for your version before depending on it.

Throttling versus the OOM killer

Kubernetes Memory QoS sequence of reclaim, throttling and OOM as container memory grows

Figure 2: What the kernel does as a Burstable container’s usage climbs past memory.low, memory.high and memory.max.

Figure 2 sequences the escalation. While usage stays below memory.low, the container is shielded from reclaim as long as unprotected cgroups still hold reclaimable pages. Between memory.low and memory.high, it competes normally, with proportional reclaim above the soft floor. Past memory.high, allocating threads are stalled and pushed into direct reclaim. Only at memory.max does the kill path open.

This is the trade the feature makes. It converts some OOM kills into slowdowns. For a batch job that is usually a gain. For a latency-sensitive service whose working set legitimately exceeds the throttle point, a slow container that never dies can be worse than a fast crash that Kubernetes restarts, because health checks may pass while tail latency collapses. The kernel documentation’s phrase, that memory.high suits setups where an external monitor relieves pressure, is the key: on Kubernetes, you are that monitor.

Tiered reservation and the node-wide limitation

With TieredReservation, the 1.36 post gives a concrete illustration: a Guaranteed pod requesting 512 MiB results in memory.min of 536870912 bytes, and the kernel will not reclaim that memory under any circumstances. A Burstable pod with the same request gets that value in memory.low instead. The 1.36 post also contrasts this with the original alpha, using an 8 GiB node whose Burstable requests total 7 GiB. Under the old behaviour all 7 GiB became hard memory.min, starving the kernel, daemons and BestEffort workloads. Under tiering, those requests become soft memory.low, so the kernel can claw back part of it to avoid a system-wide OOM.

Kubernetes Memory QoS tiered reservation protecting Guaranteed and Burstable memory on a node

Figure 3: Node memory carved into hard-protected, soft-protected and reclaimable regions under TieredReservation.

The 1.37 beta post is candid about a limitation that I consider the most important operational caveat in the whole feature. memoryReservationPolicy applies to every pod on the node. With TieredReservation, every Guaranteed pod gets memory.min and every Burstable pod gets memory.low, and there is no way to opt an individual pod in or out. A node that mixes workloads needing hard reservation with workloads that should stay reclaimable must pick one policy for all of them.

A second caveat compounds it. Hard reservation covers everything charged to the container’s cgroup, including page cache. A Guaranteed pod that reads large files can therefore hold memory the kernel would otherwise reclaim to serve its neighbours. SIG Node tracks both issues in kubernetes/kubernetes#140246, and the post invites users to describe their workloads there.

There is also a safety check on the other side. If you disable the feature after upgrading, the kubelet rejects configurations that still set memoryThrottlingFactor to anything other than the former 0.9 default, or that set TieredReservation, so you must remove those fields first. With the gate off, the kubelet resets stale protection at startup on cgroup v2 nodes: memory.min and memory.low go to zero on the root kubepods cgroup, and memory.low to zero on the Burstable cgroup. Stale container memory.high values are reset to max on paths such as restart or resize.

Observability you get with it

The 1.36 alpha introduced two kubelet metrics, both flagged ALPHA at that time: kubelet_memory_qos_node_memory_min_bytes, the total memory.min across Guaranteed pods, and kubelet_memory_qos_node_memory_low_bytes, the total memory.low across Burstable pods. If the first creeps toward the node’s physical memory, hard reservation is getting tight. I did not confirm their stability level in 1.37, so check your version’s metrics reference. Native histograms, which also graduated to beta in 1.37 according to the release announcement, are a related but separate Prometheus-side improvement; I do not cover them in depth here.

The kubelet also logs a warning at startup if the gate is enabled on a kernel older than 5.9, because memory.high throttling on older kernels can trigger a known livelock that was fixed in 5.9. The feature still works on those kernels, so the warning is informational, but it is a strong reason to treat 5.9 as your practical floor even though cgroup v2 itself needs only 5.8.

Pod-Level Resource Managers: NUMA Placement Without the All-or-Nothing Tax

Memory QoS decides how the kernel treats memory once it is allocated. Pod-Level Resource Managers address a different question: where the resources physically come from. The feature, described by Kevin Torres Martinez of Google, graduated to beta in 1.37 but, unlike Memory QoS, it is disabled by default behind the PodLevelResourceManagers feature gate. It first appeared as alpha in 1.36.

The problem it solves

The kubelet has three node-level resource managers. The CPU Manager can pin containers to exclusive cores, the Memory Manager can assign NUMA-local memory, and the Topology Manager coordinates their choices so a container’s CPUs, memory and devices sit on the same NUMA node. NUMA, non-uniform memory access, matters because a core reading memory attached to another socket pays a latency and bandwidth penalty.

Historically, getting that exclusive, aligned treatment required the pod to be Guaranteed, with integer CPU requests on every container. The announcement describes the resulting choice as all-or-nothing: give integer requests to every container in the pod, or forfeit exclusive NUMA alignment entirely. That is wasteful in modern pods, which often carry a logging agent, a telemetry exporter or a service-mesh proxy beside the real workload. Those sidecars do not need dedicated physical cores, but the rule forced you to allocate them anyway or lose alignment for the main container.

How pod-level declarations change allocation

Pod-level resources, declared under .spec.resources, let you state a budget for the whole pod rather than only per container. With PodLevelResourceManagers on, the three managers use that declaration directly in placement decisions. The practical result, in the announcement’s words, is a hybrid model: the kubelet reserves exclusive NUMA-aligned resources for the primary application containers, while non-Guaranteed sidecars are placed in a pod-isolated shared pool.

Kubernetes pod-level resource managers placing exclusive and shared containers on a NUMA node

Figure 4: Pod-level resource managers split one pod into an exclusive NUMA-aligned set and a pod-isolated shared pool.

The announcement says this gives primary workloads unthrottled, NUMA-local performance, while sidecars run with local NUMA alignment and protection from external node interference without consuming dedicated cores. I read that as a real efficiency gain for pods with many small helpers: the exclusive set stays small, and the cores you save return to the node’s shared pool for other tenants.

The beta also adds an observability change. The v1 PodResources gRPC service, PodResourcesLister, gains top-level cpu_ids and memory fields on PodResources responses. Monitoring tools and device plugins can read pod-level exclusive assignments directly, without double-counting container allocations. If you run node agents that account for pinned CPUs, such as telemetry exporters or device plugin and DRA integrations, this is the field set to move to.

Where the two features meet

Memory QoS and the resource managers are independent gates with independent defaults, but they interact through the Memory Manager and the Guaranteed QoS class. A pod that qualifies for hard memory.min reservation under TieredReservation is the same kind of pod that tends to be a candidate for NUMA-aligned memory. Stacking them gives the strongest isolation Kubernetes offers on a node: the memory is placed on the right NUMA node and the kernel is told not to take it back.

That stacking also stacks the risk. Hard-reserved, NUMA-pinned memory is capacity you cannot lend out, so a node running both on many Guaranteed pods can strand headroom. I have not seen published guidance quantifying the combined effect, so I treat it as an area for your own load tests rather than a number to quote.

One more thing worth saying plainly: the 1.37 posts do not describe a behaviour change for existing clusters from Pod-Level Resource Managers, since the gate is off. The risk is entirely opt-in, which is the opposite posture to Memory QoS, where the gate is on but the behaviour is inert until configured.

Rolling Out Memory QoS Safely: A Walk-through

The beta posture makes the rollout order matter more than the feature list. The two knobs have different blast radii, and the 1.36 post says explicitly that separating throttling from reservation lets you enable throttling first, observe behaviour and opt into reservation when the node has headroom.

Step 1: inventory your starting state

Before changing anything, confirm which kubelet configuration your nodes actually run. The upgrade effect depends on whether memoryThrottlingFactor appears in the file. Check that the node is on cgroup v2 with stat -fc %T /sys/fs/cgroup/, which prints cgroup2fs on v2 per the Kubernetes docs, and confirm the kernel is 5.9 or newer. Note which runtime and version each node uses, since container-level cgroups are the runtime’s responsibility.

Also inventory your applications’ awareness of cgroup v2. The Kubernetes docs call out specific minimums: cAdvisor v0.43.0 or later, recent OpenJDK builds such as jdk8u372, 11.0.16 and 15 and later, uber-go/automaxprocs v1.5.1 or higher, and Node.js 20.3.0 or later for reliable cgroup v2 memory-limit detection. A runtime that reads the host’s total memory instead of the pod limit will size its heap wrongly and invite OOM kills, with or without Memory QoS.

Step 2: enable throttling on a canary pool

Pick a node pool that runs tolerant, Burstable workloads such as batch or queue consumers, and set an explicit factor. A conservative choice such as 0.9, the former default, leaves the widest runway before throttling. Then watch for the signature of throttling: rising latency with flat CPU, growth in the memory.events high counter, and PSI memory pressure on the container cgroups.

The memory.events file in each cgroup records how many times the high and max boundaries were hit, and memory.pressure exposes PSI. Those are kernel facts documented in the cgroup v2 reference, and they are the most direct way to see throttling that Kubernetes-level metrics do not surface. Compare the high event rate to p99 latency per workload. If the two rise together, the factor is too aggressive for that workload.

Step 3: decide whether you need reservation at all

Tiered reservation is where the node-wide limitation bites, so make the decision deliberately. It earns its place on nodes dominated by Guaranteed workloads that must not lose pages under neighbour pressure, such as an in-memory database beside batch jobs. It is a poor fit on mixed nodes where some Guaranteed pods are large file readers, since page cache charged to a hard-protected cgroup is memory the kernel cannot take back.

Use the two kubelet metrics for capacity planning. Sum Guaranteed requests plus Burstable requests against allocatable memory. The kernel documentation warns that protecting more memory than is generally available may lead to constant OOMs, and the kubelet’s kubepods root memory.min is exactly that sum. Illustratively, on a 16 GiB node with about 14 GiB allocatable, a rule of thumb I use is to keep hard reservation to well under half of allocatable, so the kernel still has unprotected pages to reclaim before it reaches for the OOM killer. That ratio is my own heuristic, not a Kubernetes recommendation.

Step 4: roll forward by pool, with a rollback path

Roll the change out pool by pool, and keep the rollback documented. To disable Memory QoS entirely, set the gate false and remove any memoryThrottlingFactor other than 0.9 and any TieredReservation, because the kubelet will reject the configuration otherwise. On restart it clears stale memory.min and memory.low, and container memory.high values reset on restart or resize. Because those resets ride on restart paths, a rollback is not instantaneous for already-running containers.

This also interacts with in-place pod resize: a resized container has its cgroup values recomputed, so any change to requests or limits shifts the memory.high threshold with it. Rightsizing automation that nudges limits upward will quietly move the throttle point too, which is worth knowing when you correlate a latency change with a resize event.

Trade-offs, Gotchas, and What Goes Wrong

Throttling that looks like an application bug. A container held above memory.high does not crash. It stalls inside the allocation path, spending CPU in reclaim. Dashboards show moderate CPU, healthy liveness probes and rising latency. Without memory.events or PSI in your toolbox you will blame the application. If you enable memoryThrottlingFactor, add those signals before you add the config.

The quiet loss of throttling on upgrade. Teams that relied on the 1.36 alpha default of 0.9 and never wrote it into the kubelet file lose throttling when they move to 1.37. That is a safe failure, since the workload simply runs unthrottled again, but it can mask a memory-hungry pod that the old throttle was containing. Write the value explicitly if you depend on it.

Reservation that strands memory. Because the policy is node-wide, a single large Guaranteed file reader can pin page cache under memory.min. The kernel will then reclaim from unprotected neighbours or invoke the OOM killer against them. Mixed nodes should be split by taints or pools until per-pod opt-in exists. SIG Node’s tracking issue is the place to follow that.

Kernel versions. Below 5.9, memory.high throttling can trigger a livelock. The kubelet only warns. Treat the warning as a failure in a fleet with older embedded or industrial Linux images, where kernel upgrades are slow. Edge devices running vendor board-support kernels are the usual offenders, and the same caution applies to the resource interfaces described in our Azure IoT Akri overview, where device-side constraints dominate node-side tuning.

Language runtimes that ignore cgroups. Older JVMs, Node.js 18 and old automaxprocs versions misread v2 limits. A runtime that sizes its heap from host memory will push usage straight through memory.high and memory.max and defeat the whole design.

Beta means contracts can still move. Memory QoS has changed its defaults between alpha and beta, which is exactly the sort of change beta permits. The pod-level managers are beta and off by default. Pin your configuration in version control and re-read the release notes at each minor upgrade rather than assuming 1.37 semantics hold in 1.38.

What nobody has measured for you. I found no published benchmark quantifying the latency cost of memory.high throttling across workloads, and I would be suspicious of any single number. The mechanics are documented; the magnitude is workload-specific and you have to measure it on your own traffic.

Practical Recommendations

Treat the two features as separate decisions. Memory QoS is already present in 1.37 and needs no gate change, so the question is only which behaviours to turn on. Pod-Level Resource Managers requires an explicit gate and is justified only if you run NUMA-pinned, latency-critical pods with sidecars.

For most clusters, my sequence is: leave the defaults alone on day one, which the project designed to be a no-op; add memoryThrottlingFactor to a canary pool of Burstable batch workloads; instrument memory.events and PSI; and only then consider TieredReservation on nodes where Guaranteed workloads dominate. For fleets with small-memory edge nodes, be even more cautious with reservation, because a few hundred megabytes of hard-protected memory is a large share of a 4 GiB device.

A short checklist to run before enabling anything:

  • Confirm cgroup v2 on every node with stat -fc %T /sys/fs/cgroup/, and kernel 5.9 or newer.
  • Check runtimes and agents against the cgroup v2 minimums in the Kubernetes docs.
  • Write memoryThrottlingFactor explicitly if you depend on throttling, so upgrades cannot drop it.
  • Add alerting on cgroup memory.events high counts and PSI memory pressure before enabling throttling.
  • Sum Guaranteed and Burstable requests against allocatable memory before enabling TieredReservation.
  • Split mixed nodes by pool until per-pod reservation opt-in exists.
  • Keep a written rollback that removes non-default fields before disabling the gate.
  • For pod-level managers, enable the gate on one NUMA-sensitive pool and validate with the PodResources API fields.

Also glance at the rest of the release. According to the release announcement, 1.37 contains 67 enhancements: 16 stable, 23 beta and 27 alpha. Our write-up of the stable Metrics API and rootless kubelet changes covers adjacent items that touch the same kubelet configuration surface.

Frequently Asked Questions

Is Memory QoS enabled by default in Kubernetes 1.37?

The MemoryQoS feature gate is beta and on by default in 1.37, but it does nothing visible until you configure it. The default memoryThrottlingFactor is null and the default memoryReservationPolicy is None, so the kubelet writes no memory.high, memory.min or memory.low values. The project made that choice so an upgrade would not change runtime behaviour or start throttling workloads that previously ran freely.

How is memory.high calculated for a Burstable container?

The Kubernetes documentation gives memory.high = requests + memoryThrottlingFactor x (limits - requests). With a 256 MiB request, a 1 GiB limit and a factor of 0.9, the result is about 947 MiB. If a Burstable container has no limit, node allocatable memory is used in its place. Guaranteed pods get no memory.high because their requests equal their limits, and BestEffort values derive from node allocatable memory.

What is the difference between memory.min and memory.low?

memory.min is hard protection: memory within the boundary is never reclaimed, and if nothing unprotected remains, the OOM killer runs against other processes. memory.low is best-effort: memory within the boundary is spared unless unprotected cgroups have nothing left to reclaim. With TieredReservation, Kubernetes maps Guaranteed pod requests to memory.min and Burstable pod requests to memory.low, keeping hard reservation small.

Will memory.high throttling kill my pod?

No. The kernel documentation states that exceeding memory.high never invokes the OOM killer. The cgroup’s processes are throttled and pushed into heavy reclaim instead. The kill happens only at memory.max, which comes from the container’s memory limit. The risk with memory.high is slowdown rather than termination, which can hide behind healthy liveness probes while latency degrades.

Are Pod-Level Resource Managers enabled by default?

No. In 1.37 the PodLevelResourceManagers feature gate is beta and disabled by default, and it was alpha in 1.36. When enabled, the Topology, CPU and Memory Managers use pod-level .spec.resources declarations, so a pod can have exclusive NUMA-aligned resources for its primary containers while non-Guaranteed sidecars share a pod-isolated pool. The PodResources API also adds top-level cpu_ids and memory fields.

What kernel version do I need for Memory QoS?

Kubernetes documents cgroup v2 as requiring Linux kernel 5.8 or later, but recommends 5.9 or newer for Memory QoS. On older kernels, memory.high throttling can trigger a known livelock that was fixed in 5.9. The kubelet logs a startup warning below 5.9 when the gate is enabled rather than blocking the feature, so it is easy to miss unless you watch kubelet logs.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *