Kubernetes 1.37 HPA Scale-to-Zero Beta vs KEDA and Knative

Kubernetes 1.37 HPA Scale-to-Zero Beta vs KEDA and Knative

HPA Scale to Zero in Kubernetes 1.37: Beta Feature vs KEDA and Knative

For almost a decade the honest answer to “can Kubernetes scale my idle workers to zero?” was “not by itself.” The HorizontalPodAutoscaler refused to go below one replica, so every team that wanted idle queue consumers or GPU inference pods to disappear installed KEDA, Knative, or a home-grown controller. On 26 August 2026 that changed in a small but meaningful way: Kubernetes 1.37 (“Garhwal”) shipped HPA scale to zero as a beta feature, enabled by default, so a plain HorizontalPodAutoscaler with minReplicas: 0 can now park a workload and wake it again.

That matters now because idle capacity is the quiet line item in most cluster bills, especially where GPUs are involved, and “do I still need KEDA?” is suddenly a fair question. The short answer is that the native feature covers one narrow, important slice of what KEDA and Knative do, and it comes with sharp edges that can strand a workload at zero.

This post explains the mechanism, the exact requirements, the two failure modes that bite first, the cold-start arithmetic, and a decision framework for choosing between native HPA, KEDA, and Knative.

What this covers: what changed in Kubernetes 1.37, how the controller decides to go to and come back from zero, YAML for a working setup, a cold-start model, a comparison with KEDA and Knative, failure modes, and a migration checklist.

Context and Background

The HorizontalPodAutoscaler (HPA) is a control loop in kube-controller-manager. On each sync it reads a metric, compares it to a target, and patches the scale subresource of a Deployment, StatefulSet, or other scalable object. Historically the API server rejected minReplicas: 0 unless a feature gate called HPAScaleToZero was switched on, and the gate sat in alpha from Kubernetes 1.16 onward. Most managed clusters never enabled alpha gates, which is why the ecosystem filled the gap.

The incumbents each solved a different version of the problem. KEDA (Kubernetes Event-driven Autoscaling) adds a ScaledObject custom resource, an operator that polls event sources, and a metrics adapter that feeds the standard HPA for the 1-to-N range. It handles the 0-to-1 activation itself, which is why it worked on clusters where the HPA could not. Its documentation lists more than a hundred scalers for queues, databases, cloud services, and cron schedules. We covered the architecture in depth in our KEDA event-driven autoscaling architecture guide.

Knative Serving takes a request-driven view. Its autoscaler (the KnativePodAutoscaler, or KPA) scales Revisions to zero after a grace period and relies on an activator component to hold incoming HTTP requests while the first pod starts. Knative’s documented defaults include enable-scale-to-zero: true, a scale-to-zero-grace-period of 30 seconds, and a scale-to-zero-pod-retention-period of 0 seconds, according to the Knative scale-to-zero documentation.

Neither is going away. What 1.37 changes is the floor: for the common case of a queue consumer or batch worker, the platform itself can now do what previously required an add-on. The Kubernetes project’s own announcement describes the feature as API support for horizontally autoscaling workloads down to zero replicas, usable with a suitable object or external metric, and notes that before 1.37 you needed an add-on or the alpha gate (Kubernetes blog, 2 September 2026). The release itself carried 67 enhancements, of which 16 graduated to stable and 23 to beta.

One framing point before the mechanics. Scaling to zero is two problems, not one. The first is deciding that nothing is needed, which any autoscaler can do. The second is noticing demand when there are no pods to measure, and then absorbing the work that arrives during the seconds it takes to start one. Native HPA now solves the first problem and part of the second. It does not buffer anything.

How Native HPA Scale to Zero Works

Native HPA scale to zero works by letting the HPA controller set a target to zero replicas when an object or external metric shows no demand, record that decision in a ScaledToZero status condition, and scale back up when the same metric rises. CPU and memory metrics cannot trigger it, because zero pods produce no resource measurements.

HPA scale to zero control loop reading an external metric and toggling a Deployment between zero and N replicas

Figure 1: The HPA scale to zero loop. The controller keeps polling an external or object metric even when the target has zero pods, and uses the ScaledToZero condition to remember that it caused the zero state.

The figure shows the loop. A metrics source such as a queue backlog is exposed through the external metrics API (external.metrics.k8s.io) or an object metric on an existing Kubernetes object. The HPA controller reads it on every sync, which defaults to 15 seconds. When the metric falls to zero long enough to pass the downscale stabilization window, the controller sets the scale target to zero and marks the condition. When the metric rises, the same loop runs in reverse.

The feature gate and API validation

The gate is named HPAScaleToZero. According to the KEP (Kubernetes Enhancement Proposal 2021, from SIG Autoscaling), it controls two things at once: API validation that accepts minReplicas: 0, and the controller behavior that acts on it. It has existed since 1.16 and is now on by default for the first time, as reported by Techzine’s coverage of the release.

Two operational consequences follow from the gate covering both paths. If your API server and controller manager are on different versions during an upgrade, the KEP flags version skew as a risk: an API server that accepts minReplicas: 0 paired with a controller manager that does not understand it yields an HPA that validates but never acts. Upgrade the control plane as a unit and verify with a test HPA before you rely on it.

Second, downgrade is not free. The KEP’s guidance is to scale workloads back to at least one replica before disabling the gate, because a workload sitting at zero with no controller willing to wake it is exactly the stranded state you are trying to avoid.

Why only object and external metrics

A Resource metric such as CPU utilization is computed per pod as usage divided by requests, then averaged. With zero pods the numerator and denominator both disappear. The KEP lists resource-metric scale to zero as an explicit non-goal. Pods metrics have the same flaw, since they are also sourced from running pods.

Object metrics describe a Kubernetes object, for example the request rate reported for an Ingress. External metrics describe something outside the cluster, such as the depth of an SQS queue or a Pub/Sub backlog, and are served by an adapter that implements the external metrics API. Both exist independently of the pods being scaled, which is the property that makes waking from zero possible. In practice this means you need a metrics adapter in the cluster: the Prometheus adapter, a cloud provider adapter, or KEDA’s own metrics server, which can run alongside native HPAs.

The ScaledToZero condition

The subtlest part of the design is the ScaledToZero condition on the HPA status. The KEP introduces it to distinguish a workload the HPA scaled to zero from one a human or another controller paused at zero. When the HPA moves 1 to 0 it records ScaledToZero=True, and when it moves 0 to 1 it checks for that condition first. Techzine reports the condition flips to False when scaling up.

The reasoning is sound. A Deployment with replicas: 0 is a common way to say “off on purpose,” used in maintenance, incident response, and blue-green cutovers. If the HPA woke every such workload whenever a metric moved, it would override operators. The condition makes the HPA touch only what it parked. It also creates the first sharp edge, covered below: a workload that reaches zero by any other route will not be woken.

What it deliberately does not do

The KEP lists request buffering at the service level as a non-goal. Nothing in native HPA sits in the request path. If an HTTP request hits a Service with zero endpoints, the connection fails or times out; no component holds it until a pod is ready. That is the single biggest functional gap against Knative and the KEDA HTTP add-on, and it is why the feature targets queue-driven, batch, and scheduled workloads rather than synchronous APIs.

Walk-through: A Working Setup and the Cold-Start Budget

The fastest way to understand the feature is to build the smallest useful configuration and then time what happens between the first message and the first processed job.

A queue worker that scales to zero

The example below scales a worker Deployment on an external metric for queue depth. The metric name and selector depend on your adapter; treat them as placeholders. The structure is what matters: minReplicas: 0, and an External metric with an AverageValue target.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: invoice-worker
spec:
  replicas: 1            # start at 1, never 0, so the HPA owns the first scale-down
  selector:
    matchLabels: { app: invoice-worker }
  template:
    metadata:
      labels: { app: invoice-worker }
    spec:
      containers:
        - name: worker
          image: registry.example.com/invoice-worker:2.4.1
          resources:
            requests: { cpu: 250m, memory: 256Mi }
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: invoice-worker
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: invoice-worker
  minReplicas: 0
  maxReplicas: 20
  metrics:
    - type: External
      external:
        metric:
          name: queue_messages_ready
          selector:
            matchLabels: { queue: invoices }
        target:
          type: AverageValue
          averageValue: "10"
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300
    scaleUp:
      stabilizationWindowSeconds: 0

Note the Deployment starts at one replica on purpose. As the next section explains, a workload that is applied at zero has no ScaledToZero condition, so the HPA treats it as intentionally paused. Let the controller perform the first 1-to-0 transition itself.

With AverageValue: "10", the HPA targets ten ready messages per replica. A backlog of 95 messages gives ceil(95 / 10) = 10 replicas, capped by maxReplicas. A backlog of zero gives zero replicas. This arithmetic is the standard HPA formula, desired = ceil(current replicas * current metric / target metric), except that for an external AverageValue metric the controller divides the total metric value by the target instead, since the metric is not per-pod.

Three details deserve attention. The scaleDown.stabilizationWindowSeconds of 300 is the documented HPA default; it makes the controller act on the highest recommendation over the last five minutes, which damps flapping but also means an idle queue holds capacity for at least that long. Setting scaleUp.stabilizationWindowSeconds: 0 is also the default and is shown for clarity. And the whole loop only runs as fast as the weakest link in the chain, which brings us to latency.

Where the seconds go

Sequence diagram of HPA scale to zero cold start from first message to worker ready

Figure 2: Cold-start sequence for native HPA scale to zero. Every stage between the first enqueue and the first processed job adds latency, and none of them is buffered by the HPA itself.

Figure 2 lays out the path. The producer enqueues the first message. The metrics adapter has to observe it, then the HPA controller has to read the adapter, decide, and patch the scale subresource. Then the scheduler places the pod, the kubelet pulls the image if it is not cached, the container starts, and the readiness probe passes. Only then does a worker consume the message.

The numbers below are illustrative, not measured. They show how the stages compose, and your own cluster will differ.

Stage Optimistic Typical Pessimistic
Adapter scrape interval 5 s 15 s 30 s
HPA sync period 2 s 15 s 15 s
Scheduling on existing node 1 s 2 s 5 s
Node provisioning, if none free 0 s 0 s 60 to 120 s
Image pull, cached vs cold 0 s 5 s 60 s
Container start and readiness 2 s 10 s 30 s
Total, worst case per column 10 s 47 s about 5 min

Summing the worst case per stage, the “typical” column gives about 47 seconds and the pessimistic column approaches five minutes once node provisioning enters. Averages are kinder than worst cases, because the scrape and sync loops are not synchronized with the enqueue and on average add half their period. The point is that cold-start latency is dominated by a few terms you can attack: the adapter scrape interval, the image pull, and node availability.

For a batch pipeline that tolerates a minute of delay, this is irrelevant. For a payment callback queue with a 10-second SLA it is disqualifying, and you would keep minReplicas: 1 and accept the idle cost.

Attacking the cold start

Four levers matter, in rough order of payoff. First, shorten the adapter’s collection interval, remembering that every tightening costs load on the metric source. Second, keep images small and pre-pulled; on nodes that run the workload regularly the image is usually cached, and a DaemonSet that pre-pulls the image onto a pool removes the 60-second term. Third, make readiness cheap: a worker that loads a 2 GB model at startup has a cold start measured in minutes whatever the autoscaler does. Fourth, keep a warm node: if the cluster autoscaler or Karpenter has to provision a node, that term dominates everything else.

That last point connects to node-level scaling. Pod scale to zero without node scale-down saves nothing, and node scale-down with an aggressive consolidation policy turns every wake-up into a node launch. Our Karpenter node autoscaling deep dive covers how consolidation timing interacts with workloads that appear and vanish, and the same trade-off appears in the Karpenter versus Cluster Autoscaler GPU comparison, where a GPU node launch can dwarf every pod-level delay.

What the saving actually looks like

Consider a worker pool that needs four replicas of 2 vCPU and 8 GiB during a two-hour nightly batch and is idle the other 22 hours. At minReplicas: 1, one replica runs around the clock. The saving from scaling to zero is the idle replica’s 22 hours, which is a fraction of one replica’s cost, not of the pool’s. The arithmetic is illustrative: one 2 vCPU replica for 22 hours is 44 vCPU-hours a day. If the node it occupies would otherwise be released, that is real money; if the node stays up for other tenants, the saving is zero.

This is the part vendors rarely say plainly. Scaling pods to zero only turns into a bill reduction when the nodes underneath can also disappear, and when the idle replica count was non-trivial to begin with. With a GPU worker the maths flips: one idle GPU replica running 22 hours a day is a large number, and scale to zero pays for itself immediately provided your node pool can release the instance.

Native HPA vs KEDA vs Knative: A Decision Framework

Having seen the mechanism, the comparison is mostly about what sits around the 0-to-1 transition. All three ultimately rely on an HPA or an HPA-like loop for the 1-to-N range; they differ in how they detect demand at zero, whether they hold requests, and how much machinery you operate.

Decision flow comparing native HPA scale to zero with KEDA and Knative by workload type

Figure 3: Choosing between native HPA scale to zero, KEDA, and Knative by workload type. The deciding questions are the number of event sources and whether requests must be buffered during cold start.

How KEDA handles zero differently

KEDA splits the problem into an activation phase and a scaling phase. According to the KEDA scaling documentation, whether a workload moves between zero and one is decided by each scaler’s IsActive result, while the metrics adapter serves values to an HPA that KEDA creates for the 1-to-N range. Each trigger has an optional activationThreshold that defaults to 0, and the docs state that activation happens only when the value is strictly greater than the threshold. With defaults, one queued message activates the workload.

The practical difference from native HPA is that KEDA’s own operator polls the event source directly to decide activation. It does not need a metric adapter’s value to pass through the HPA’s stabilization logic for the 0-to-1 step. KEDA also exposes an annotation, autoscaling.keda.sh/force-activation: "true", to force activation by hand. You can use it for maintenance windows or warm-ups, and the native feature has no equivalent.

The other differences are breadth and decoupling. Scalers for Kafka, RabbitMQ, Azure Service Bus, AWS SQS, Prometheus, cron, and a hundred more are maintained as part of the project. With native HPA you bring your own adapter and are responsible for exposing a correctly shaped external metric. KEDA’s ScaledJob also runs one Kubernetes Job per work item, a pattern the HPA has no analogue for. For release-level changes and security fixes, see our KEDA 2.21 migration notes.

How Knative handles zero differently

Knative’s design is request-first. The autoscaler counts concurrent requests or requests per second, scales a Revision to zero after the grace period, and places an activator in the data path while the Revision has no pods. The activator receives requests, signals the autoscaler to scale up, and forwards the request once an endpoint is ready, so the caller sees added latency rather than an error. That is the capability the KEP excludes from native HPA by design.

The cost is a larger platform. Knative Serving brings its own CRDs (Service, Route, Revision, Configuration), a networking layer, and a different deployment model from a plain Deployment. Its scale-to-zero is global in one respect that surprises people: the documentation states enable-scale-to-zero is configured cluster-wide rather than per Revision, and scale to zero requires the KPA rather than the HPA class.

Side-by-side

Dimension Native HPA 1.37 beta KEDA Knative Serving
Status Beta, on by default Mature CNCF project Mature CNCF project
Wakes from zero on Object or external metric via HPA Scaler IsActive polled by operator Request arrival at activator
Request buffering None None, unless HTTP add-on used Yes, via activator
Event sources Whatever your adapter exposes 100+ built-in scalers HTTP and gRPC focus
Extra components Metrics adapter Operator, metrics server, admission webhook Controller, autoscaler, activator, networking layer
Workload type Any scalable resource Deployments, StatefulSets, custom resources, Jobs Knative Services only
Failure when metric breaks HPA holds or stays at zero Operator reports unhealthy scaler Requests queue at activator
Operational surface Smallest Medium Largest

The matrix points to a clean rule. If you already run a metrics adapter and your workload is a queue consumer, native HPA is enough and removes a component. If you have many different triggers, need ScaledJob, or want activation decoupled from HPA stabilization, KEDA still earns its place. If you serve synchronous HTTP traffic and users must not see errors on the first request, you need a request buffer, which means Knative or the KEDA HTTP add-on.

Running native HPA and KEDA together

You do not have to choose one for the whole cluster. KEDA’s metrics server exposes external metrics that a hand-written HPA can consume, so a team can use KEDA purely as an adapter library for the sources it supports while leaving the scaling decision to a native HPA. The one rule that must not be broken is that a single scale target should have exactly one autoscaler. A ScaledObject creates and owns an HPA for its target; pointing a second HPA at the same Deployment produces two controllers fighting over the scale subresource, and the symptoms (replica counts oscillating, conditions reporting conflicts) are miserable to debug.

A reasonable migration pattern follows from that. Keep KEDA for workloads that use its activation threshold, ScaledJob, or cron triggers. Move simple queue consumers to native HPA one at a time, in a non-production namespace first, and delete the ScaledObject before applying the HPA so ownership is never ambiguous. Treat the move as a simplification project, not a performance project: the native path is not faster, and its behavior at the 0-to-1 boundary is tied to the HPA’s loop timing rather than KEDA’s own polling.

Trade-offs, Gotchas, and What Goes Wrong

Beta means the feature is on by default and the API is considered stable enough to build on, but the project still expects feedback and may refine behavior. The failure modes below are the ones that follow directly from the design, and the first two are the situations that strand a workload at zero. They are derived from the KEP’s design and from third-party write-ups of the release, not from measurements I took.

Troubleshooting flow for a workload stuck at zero replicas under HPA scale to zero

Figure 4: Diagnosing a workload stuck at zero. Check the metric type first, then the ScaledToZero condition, then whether the adapter returns a real value.

Deadlock one: resource metrics at zero

If the HPA lists only CPU or memory, minReplicas: 0 is the wrong configuration. With no pods there is no utilization, so there is nothing to wake on. The fix is to include at least one object or external metric. Be careful with mixed metric lists: when several metrics are configured, the HPA computes a replica count for each and takes the largest. A CPU metric alongside a queue metric is legal for the 1-to-N range, but the CPU metric contributes nothing at zero, so the queue metric alone must carry the wake-up.

Deadlock two: zero without the condition

The ScaledToZero condition is what permits waking. A workload that reaches zero in any other way, such as being applied with replicas: 0, scaled down with kubectl scale during an incident, or reaped by a platform idle-cleanup job that edits spec.replicas, carries no such condition. The HPA reads that as a deliberate pause and leaves it alone.

This produces an ugly operational pattern. An engineer scales a worker to zero to stop a bad consumer, fixes the bug, and expects the HPA to resume. It does not, because the HPA never took the workload to zero. The remedy is procedural: deploy at one replica, route every replica change through the HPA’s inputs, and teach the on-call runbook that kubectl scale --replicas=1 is the way to restart a manually paused autoscaled workload. If your GitOps tool renders replicas: 0 into the manifest, remove the field and let the HPA own it, which is also the recommended practice for any autoscaled Deployment.

The metric that returns nothing

External metrics depend on an adapter that must answer. If the adapter pod crashes, its APIService becomes unavailable, or the underlying system (Prometheus, a cloud API) is unreachable, the HPA sees an error or an unknown value. The standard HPA behavior on missing metrics is conservative: it does not scale down on missing data, and it can refuse to scale up on partial data. At zero replicas, an unreadable metric means the workload stays at zero while messages accumulate.

Monitor this directly. Alert on the HPA’s ScalingActive condition going false, and alert on queue age independently of the autoscaler, so a silent adapter failure shows up as latency in your SLO rather than as a surprise. KEDA has the same dependency conceptually, but its operator surfaces unhealthy scalers as resource status, which makes the failure easier to see.

Fast queues and the stabilization window

The default scale-down stabilization of 300 seconds protects against flapping, but a bursty queue interacts with it in awkward ways. A producer that emits one message every four minutes will keep the workload warm forever if you scale on “any message present,” because the window never sees a long enough quiet period. Conversely, shrinking the window to save money makes each message pay a cold start. Choose the window from the traffic’s inter-arrival distribution, not from a default.

Version skew, downgrades, and the upgrade path

The KEP calls out version skew between kube-apiserver and kube-controller-manager as a risk, and calls out rollback as a source of stranded workloads. The upgrade checklist is therefore specific: upgrade the control plane together, verify a canary HPA scales to zero and back, and before any downgrade scale all minReplicas: 0 workloads to one replica. Managed Kubernetes services also trail upstream, so confirm the version your provider runs before assuming the gate is on. I could not verify per-provider availability dates for 1.37 as of this writing.

What native HPA will not give you

No request buffering, no per-trigger activation thresholds, no ScaledJob equivalent, no scheduled scaling. There is also no scale-to-zero for workloads exposed to synchronous traffic unless something in front holds requests. Teams that read “HPA scale to zero” and apply it to an HTTP service behind an Ingress will see 502 or 503 responses on the first request after idle. That is not a bug; it is the documented non-goal.

Practical Recommendations

Start by classifying workloads by who is waiting. If nobody is waiting synchronously, for example a nightly reconciliation job, an email sender, or a media transcoder fed from a queue, native HPA scale to zero is the simplest option and you should adopt it where you already have a metric adapter. If a human or an upstream service is waiting on the response, keep a warm replica or use a request-buffering layer.

For GPU workloads the calculus favors scaling to zero most strongly, but the cold-start budget is also the worst: node launch, driver readiness, and model load can total several minutes. Pair pod-level scale to zero with a node pool that releases GPU instances and a model-loading path that is as fast as you can make it. Quantify the idle cost first, using your own billing data, before committing to an architecture that adds minutes to every cold request.

For existing KEDA users there is no reason to rush. KEDA keeps working, it is mature, and it does things the beta feature does not. Migrate only the workloads whose KEDA configuration is trivial, and do it for the operational simplicity of one fewer component.

A short checklist for rollout:

  • Confirm the cluster runs Kubernetes 1.37 or later and the HPAScaleToZero gate is enabled on both API server and controller manager.
  • Install and health-check a metrics adapter that serves the external or object metric you need.
  • Deploy the workload at one replica and let the HPA perform the first scale-down.
  • Include no CPU-only configuration for minReplicas: 0; always include an object or external metric.
  • Set the downscale stabilization window from measured inter-arrival times.
  • Measure cold start end to end and compare it with the SLO of the slowest consumer.
  • Alert on queue age, on HPA ScalingActive, and on ScaledToZero staying true while the queue is non-empty.
  • Ensure nodes can be released, otherwise the saving is only on paper.
  • Write the on-call rule: restart a paused workload with kubectl scale --replicas=1, never by editing manifests.
  • Before any downgrade, scale all zero-replica workloads to one.

Frequently Asked Questions

Is HPA scale to zero enabled by default in Kubernetes 1.37?

Yes. The HPAScaleToZero feature gate entered beta in Kubernetes 1.37 and is enabled by default for the first time, after sitting in alpha since 1.16. You then enable behavior per workload by setting minReplicas: 0 on an HPA that uses at least one object or external metric. Managed services may lag the upstream release, so confirm your provider’s Kubernetes version and gate status before relying on it in production.

Can the HPA scale to zero on CPU or memory?

No. CPU and memory are resource metrics computed from running pods, so at zero replicas there is nothing to measure and nothing to trigger a wake-up. The feature requires an object or external metric, such as a queue backlog or a request rate from an existing object. You can still list CPU alongside such a metric to govern the 1-to-N range, but the non-resource metric must carry the 0-to-1 decision.

Does native scale to zero replace KEDA?

For simple queue-driven workloads where you already run a metrics adapter, it can. It does not replace KEDA’s breadth of built-in scalers, per-trigger activation thresholds, ScaledJob, or the force-activation annotation. Many teams will use both: KEDA for complex or multi-source triggers and native HPA where a single external metric suffices. Whichever you choose, give each scale target exactly one autoscaler to avoid controllers fighting over replicas.

Why did my HPA not wake a workload at zero?

The most common cause is that the HPA never scaled it to zero. The ScaledToZero condition is set only when the HPA performs the 1-to-0 transition, and the HPA will not wake a workload that reached zero through replicas: 0 in a manifest, a manual kubectl scale, or another controller. Scale it to one replica manually, let the HPA own the next cycle, and also check that the external metric returns a real value.

What happens to HTTP requests while a pod is starting?

With native HPA, nothing holds them. If a Service has no ready endpoints the request fails or times out, so synchronous APIs behind an Ingress will return errors on the first request after idle. Knative’s activator buffers requests during cold start, and the KEDA HTTP add-on is designed for a similar role. For plain Kubernetes, either keep one warm replica or put a buffering layer in front.

Does scaling pods to zero reduce my cloud bill?

Only if the nodes underneath can also be released. Pods at zero free requests, but a node that stays up for other tenants or that your node autoscaler keeps for consolidation delay costs the same. The savings are largest for expensive nodes such as GPU instances and for pools dedicated to the workload. Verify by comparing node-hours before and after, not pod-hours, since pod-hours are not what you are billed for.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *