KServe vs Ray Serve vs BentoML: Model Serving on Kubernetes ADR

KServe vs Ray Serve vs BentoML: Model Serving on Kubernetes ADR

KServe vs Ray Serve vs BentoML: Model Serving on Kubernetes ADR

Most teams pick a model serving platform the way they pick a text editor: by whatever the first engineer on the project already knew. Then eighteen months later, with fourteen models, three GPU types and an LLM workload nobody planned for, they discover that the choice quietly decided who owns autoscaling, how canaries work, whether scale-to-zero is realistic, and which team gets paged when a GPU node drains. The three serious open-source options, KServe, Ray Serve and BentoML, make genuinely different bets about where those responsibilities should live.

This matters more in 2026 because the center of gravity moved. KServe shipped a production-grade LLM resource in March, Ray Serve grew a dedicated LLM layer with prefill/decode disaggregation, and BentoML kept pushing a Python-first developer loop. A KServe vs Ray Serve decision is no longer about sklearn latency; it is about who runs your inference fleet.

This is written as an architecture decision record. You will leave with a mechanism-level comparison, a decision matrix, concrete failure modes, and a recommendation rule you can defend in a design review.

What this covers: the control-plane model of each platform, autoscaling and scale-to-zero mechanics, GPU and LLM support, canary and rollout behavior, operational cost, a decision matrix, trade-offs and gotchas, and a checklist for choosing.

Context and Background

Model serving on Kubernetes splits into two layers that people constantly conflate. The first is the inference runtime: the process that loads weights and executes forward passes, such as vLLM, Triton, TorchServe, ONNX Runtime or a plain Python class. The second is the serving platform: the control plane that deploys that runtime, routes requests to it, scales it, rolls out new versions, and exposes a stable API. KServe, Ray Serve and BentoML are platforms. They differ mainly in how much of the Kubernetes machinery they expose and how much they hide.

KServe began inside the Kubeflow project as KFServing and became a standalone project. According to the CNCF, KServe became an incubating project in November 2025, and its repository now describes itself as a standardized distributed generative and predictive AI inference platform for Kubernetes. Its identity is Kubernetes-native: everything is a custom resource, reconciled by controllers, composable with Knative, Istio, Gateway API and KEDA.

Ray Serve comes from a different lineage. It is the serving library of Ray, a distributed Python runtime that also hosts training, data processing and reinforcement learning. On Kubernetes it is deployed through the KubeRay operator and its RayService custom resource. Ray Serve’s identity is programmable: you compose deployments in Python, and the Ray cluster is the scheduler.

BentoML starts from the developer. You write a Python class decorated with @bentoml.service, build it into a versioned artifact called a Bento, and ship it to a container registry. From there you can deploy to your own cluster or to BentoCloud, the vendor’s managed platform. BentoML’s identity is packaging and developer velocity.

Those three identities predict most of what follows. If you already run Kubernetes platform engineering as a discipline, KServe fits your existing mental model. If you already run Ray for training or batch inference, Ray Serve removes a second scheduler. If your data scientists need to ship an API without learning CRDs, BentoML shortens the path.

For the LLM-specific end of this landscape, see our deep dive on llm-d and prefill/decode disaggregation on Kubernetes, which is now directly relevant to KServe, and our comparison of vLLM, SGLang and TensorRT-LLM, the runtimes all three platforms can wrap. The KServe project’s own documentation is at kserve.github.io.

A note on evidence. I did not run a fresh head-to-head benchmark for this article, and I do not publish throughput numbers I did not measure. Where I cite figures, they come from the vendors’ own documentation and are labelled as such. Treat any third-party “X is 3x faster than Y” claim for this category with suspicion: the runtime, batch policy, hardware and model dominate the result, not the platform wrapper.

Three Control Planes, Three Philosophies

KServe, Ray Serve and BentoML differ in who owns the control plane. KServe delegates it to Kubernetes controllers and CRDs, Ray Serve delegates it to a Ray cluster with its own scheduler and autoscaler, and BentoML delegates it to a build-and-deploy workflow that targets either Kubernetes or BentoCloud.

KServe vs Ray Serve vs BentoML control planes on Kubernetes

Figure 1: The three platforms place their control plane in different layers, but all end up as model pods on GPU nodes.

Figure 1 shows the shared destination and the divergent paths. All three ultimately run a container with a model on a node. What differs is the number of controllers between your YAML and that container, and which of them you must debug at 3 a.m.

KServe: the model is a Kubernetes resource

In KServe you declare an InferenceService for predictive models, or, since version 0.17, an LLMInferenceService for generative ones. A controller turns that into Deployments, Services, autoscalers and routing. There are two deployment modes for classic InferenceService: a Knative-based serverless mode and a RawDeployment mode that uses plain Kubernetes Deployments and the Horizontal Pod Autoscaler.

The design argument is composability. Because everything is a CRD, your existing GitOps pipeline, admission policies, network policies, quotas and observability stack apply unchanged. A platform team can offer data scientists a small InferenceService manifest and keep the rest of the machinery private. The cost is that you operate real infrastructure: the KServe controllers, a networking layer, optionally Knative, and for LLMs the Gateway API plus the Gateway API Inference Extension.

KServe also standardizes the wire protocol. Its Open Inference Protocol (the v2 protocol, shared with Triton and others) gives you a consistent REST and gRPC surface across frameworks, which matters when downstream services should not care which runtime serves a given model. For generative models it exposes OpenAI-compatible endpoints through the vLLM-based runtime.

Ray Serve: the model is a Python deployment in a cluster

In Ray Serve you write deployments, which are Python classes or functions with replica counts and resource requests, and compose them into an application graph. On Kubernetes, a RayService resource from KubeRay describes the Ray cluster and the Serve config together. According to the Ray documentation, KubeRay supports upgrades without downtime and head node fault tolerance, and the application is described through the serveConfigV2 field.

The design argument is programmability. A single Serve application can hold a preprocessing deployment, three model deployments, a business-logic router and a postprocessor, all calling each other through handles with no network hop through an ingress. Fractional GPU requests, model multiplexing and dynamic batching are first-class. The cost is a second scheduler: Ray places actors on Ray nodes, and KubeRay places Ray nodes on Kubernetes. When something is pending, you must know which layer it is pending in.

BentoML: the model is a versioned artifact

BentoML’s unit is the Bento: code, dependencies, model references and configuration packaged reproducibly. A service declares its resources and traffic policy in the decorator. Deployment then happens through BentoCloud, or through a self-managed operator path on Kubernetes. The design argument is that the artifact, not the cluster, is the product. A data scientist can run the same service locally with a single command and get the same behavior in production.

The cost is that the richest operational features, notably the autoscaling policies and scale-to-zero with an external request queue documented by BentoML, are tied to BentoCloud’s control plane. If you self-host, you assemble more of that yourself or accept a narrower feature set. Verify exactly which features your chosen deployment path includes before committing, because documentation for the managed platform does not always map one-to-one to a self-hosted install.

The thesis: choose the owner of your queue

Here is the angle I think most comparisons miss. The most consequential difference is not features; it is who owns the request queue and the concurrency limit. Autoscaling quality is bounded by how accurately the scaler knows how much work is waiting. In KServe that signal comes from the Knative queue-proxy, from HPA metrics, or for LLMs from the endpoint picker and Prometheus metrics. In Ray Serve it comes from the Serve proxy and replica queue lengths. In BentoML it comes from the declared concurrency and an optional external queue.

If you can state your queue owner and your concurrency number, you can predict the platform’s autoscaling behavior. If you cannot, no platform will rescue you.

Deeper Analysis: Autoscaling, GPUs, LLMs and Rollouts

Autoscaling mechanics compared

Autoscaling is where the three platforms feel most different in production, so it deserves mechanism-level treatment rather than a feature checkbox.

Request queue ownership and autoscaling signals in KServe, Ray Serve and BentoML

Figure 2: Each platform measures load at a different point, which determines how fast and how accurately it scales.

KServe with Knative (serverless mode). The Knative Pod Autoscaler scales on concurrency by default, with requests per second, CPU and memory as alternatives. According to KServe’s documentation, the autoscaler uses a 60-second stable window and can enter panic mode, switching to a 6-second window when observed concurrency reaches twice the target. Setting minReplicas: 0 enables scale-to-zero, which the documentation highlights as valuable for GPU workloads. The same documentation notes the Knative Pod Autoscaler is only supported in Knative mode, not in RawDeployment. In raw mode you use HPA, which scales on CPU, memory or custom metrics and cannot scale to zero natively; teams add KEDA for that.

The practical consequence: Knative gives you request-aware scaling and scale-to-zero, at the price of an extra component (Knative Serving plus a networking layer) and a queue-proxy sidecar on every pod. For small CPU models that overhead is trivial. For a 70B-parameter model that takes minutes to load, scale-to-zero is a trap, which I return to in the gotchas section.

Ray Serve. The Ray documentation defines the central knobs. target_ongoing_requests is the average number of in-flight requests per replica the autoscaler tries to maintain; its default became 2 in Ray 2.32. max_ongoing_requests is the per-replica queue limit the proxies respect, and the docs recommend keeping it roughly 20 to 50 percent above the target. min_replicas defaults to 1, max_replicas must exceed it, and upscale_delay_s defaults to 30 seconds. The docs warn that setting min_replicas to 0 saves cost but raises tail latency.

There is an important two-level structure here that is easy to miss. Serve autoscales replicas inside a Ray cluster. KubeRay, in turn, can autoscale Ray worker pods so the cluster has room for those replicas. A burst therefore passes through two scalers in series: first Serve decides it wants more replicas, then Ray finds no free GPU, then the Ray autoscaler requests more nodes, then Kubernetes provisions pods, and possibly the cluster autoscaler adds machines. Each stage adds delay, and the stages are tuned independently.

BentoML on BentoCloud. BentoML’s documentation defines concurrency as the number of simultaneous requests a service can process, set in the decorator, for example traffic={"concurrency": 32}. The autoscaler computes desired replicas as the ceiling of current concurrent requests divided by the concurrency value; the docs’ own example is 100 concurrent requests at concurrency 32 yielding 4 replicas. Without a concurrency setting, scaling falls back to CPU utilization only. Setting min_replicas: 0 enables scale-to-zero, and an optional external queue holds requests while replicas start, with the documented trade-off of added latency in exchange for protection against overload. Stabilization windows between 0 and 3,600 seconds can be configured to damp flapping.

The formula is attractive because it is explainable in one sentence. Its weakness is that it treats every request as equal cost, which is false for LLMs, where one request may produce 20 tokens and another 4,000.

GPU handling and sharing

None of the three platforms invents GPU sharing; all of them lean on Kubernetes and the NVIDIA stack. The differences are in how easily you can express fractional or multi-GPU placement.

KServe passes standard nvidia.com/gpu resource requests, and node selectors, tolerations and device-plugin features such as MIG slices or time-slicing work because they are ordinary Kubernetes. If you want to understand when those mechanisms help, our guide to GPU sharing with MIG, time-slicing and MPS on Kubernetes applies unchanged.

Ray Serve adds its own resource model on top: a deployment can request num_gpus: 0.5, and Ray packs two replicas onto one GPU by bookkeeping, not isolation. That is convenient and dangerous in equal measure, because Ray does not enforce memory limits between fractional tenants. Two replicas that each believe they own half the VRAM can still out-of-memory each other. The Ray LLM documentation lists fractional GPU serving as a supported pattern, so the capability is real; the isolation guarantee is yours to engineer.

BentoML expresses resources in the service decorator, for example the GPU count and type, and relies on the underlying scheduler. On BentoCloud you select instance types; self-hosted, you again inherit Kubernetes semantics.

LLM serving: the part that changed in 2026

Generic model serving is a mostly solved problem. LLM serving is not, because the unit of work is a stream of tokens, the state is a large KV cache, and the cost profile splits into a compute-bound prefill phase and a memory-bandwidth-bound decode phase. A platform that treats an LLM as just another HTTP backend will route badly and scale badly.

Prefill decode disaggregation and cache aware routing for KServe LLM inference

Figure 3: Cache-aware routing plus separate prefill and decode pools, the pattern KServe v0.17 and Ray Serve LLM both support.

KServe LLM inference. According to the KServe v0.17 release announcement, published for the March 13, 2026 release, LLMInferenceService moved from experimental to production-ready. It is described as a CRD for generative workloads supporting distributed inference across multiple nodes and GPUs, with tensor, data and expert parallelism. The release integrates the Gateway API Inference Extension (version 1.3.0 per the announcement), where an Endpoint Picker routes requests by cached prefix blocks instead of round-robin, aiming to reduce time to first token in multi-turn conversations. It also supports disaggregated prefill and decode using the NIXL connector over RDMA, and a Workload Variant Autoscaler that works with KEDA or HPA while watching KV cache utilization and queue depth through Prometheus. Envoy AI Gateway integration provides token-based rate limiting rather than request counting.

Other 0.17 details from the same announcement are worth noting for operators: the project split into three independently installable components (kserve, llmisvc and the optional localmodel caching), the Helm chart was restructured from one chart into ten, and upgrading from v0.16 cannot be done with a plain helm upgrade. The announcement lists vLLM 0.15.1 and Gateway API 1.4.0 among the bundled versions. Version numbers move quickly; check the current release before you pin anything. KServe and the llm-d project also publish a joint integration story, described in the KServe and llm-d blog post.

Ray Serve LLM. Ray’s ray.serve.llm module specializes Serve primitives for LLMs. Its documentation lists serving patterns for small, medium and large models, vision and reasoning models, multi-LoRA deployment, cross-node parallelism, data-parallel attention, fractional GPU serving, prefill/decode disaggregation, KV cache offloading and prefix-aware routing, with vLLM and SGLang integrations. Because Serve is already a Python programming model, expressing a custom router or a mixed pipeline (embedding model, reranker, LLM) is natural. One caution from the public issue tracker: disaggregation support depends on the engine version and KV-transfer backend, and there is a request for SGLang disaggregation support still open at the time of writing. Read the current docs for your Ray and vLLM versions before relying on it.

BentoML. BentoML supports LLM serving by wrapping engines such as vLLM in a service class, and its OpenLLM project targets running open models. The platform’s concurrency formula and adaptive batching were designed with request-level workloads in mind. For LLMs the engine’s own continuous batching does the work, as explained in our piece on continuous batching for LLM inference. I did not find BentoML documentation describing built-in prefix-cache-aware routing or native prefill/decode disaggregation; treat it as not documented rather than absent, and verify with the vendor.

Code: the same model, three declarations

To make the ergonomics concrete, here is a minimal text classifier in each platform. These are condensed sketches that illustrate shape, not production configs.

# KServe: declarative, Kubernetes-native
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: sentiment
spec:
  predictor:
    minReplicas: 0
    canaryTrafficPercent: 10
    model:
      modelFormat:
        name: huggingface
      storageUri: s3://models/sentiment/v2
      resources:
        limits:
          nvidia.com/gpu: "1"
# Ray Serve: programmable, composed in Python
from ray import serve

@serve.deployment(
    autoscaling_config={
        "min_replicas": 1,
        "max_replicas": 8,
        "target_ongoing_requests": 2,
    },
    ray_actor_options={"num_gpus": 1},
)
class Sentiment:
    def __init__(self):
        from transformers import pipeline
        self.pipe = pipeline("sentiment-analysis", device=0)

    async def __call__(self, request):
        return self.pipe((await request.json())["text"])

app = Sentiment.bind()
# BentoML: service class plus traffic policy
import bentoml
from typing import List

@bentoml.service(
    resources={"gpu": 1},
    traffic={"concurrency": 32, "timeout": 30},
)
class Sentiment:
    def __init__(self):
        from transformers import pipeline
        self.pipe = pipeline("sentiment-analysis", device=0)

    @bentoml.api(batchable=True, max_batch_size=32, max_latency_ms=100)
    def classify(self, texts: List[str]) -> List[dict]:
        return self.pipe(texts)

The shapes tell the story. The KServe version has no Python and one resource; the platform team can template it. The Ray version has no YAML for the application logic, and the autoscaling policy sits next to the code that it governs. The BentoML version is the shortest path from a notebook, and adaptive batching is a decorator argument. BentoML’s documentation describes the batch window this way: the system groups requests until either the size limit or the latency limit is reached, adjusting to traffic, so a max_latency_ms of 100 bounds the extra wait you accept for batching.

Canaries, rollouts and versioning

Rollout behavior is the second place the platforms diverge sharply.

KServe makes canaries a field. In Knative mode, canaryTrafficPercent splits traffic between the latest ready revision and the previous one; promoting means removing the field or setting it to 100. Because revisions are Knative objects, rollback is a single traffic edit. In RawDeployment mode you lose Knative revisions, so canary behavior depends on your ingress or service mesh, and you will likely lean on Gateway API traffic splitting or Argo Rollouts.

Ray Serve rolls out at the application level. KubeRay’s RayService supports upgrades with a standby cluster so traffic switches only after the new cluster is healthy; the Ray documentation describes zero-downtime and incremental upgrade modes. The consequence is that a rollout may briefly double your GPU footprint, since two Ray clusters coexist. For a fleet of eight-GPU nodes that is a real capacity planning input, not a footnote.

BentoML versions the artifact. Each Bento has a tag, and a deployment references a specific one, so rollback means redeploying the earlier tag. Traffic splitting between versions is a deployment-platform feature rather than something intrinsic to the Bento format; confirm its availability in your chosen target.

The decision matrix

The matrix below compresses the comparison. Ratings are my qualitative judgment from the documentation cited above, not measured results, and they assume a team operating on Kubernetes with GPUs.

Dimension KServe Ray Serve BentoML
Primary abstraction Kubernetes CRD (InferenceService, LLMInferenceService) Python deployment graph on a Ray cluster Versioned Bento artifact
Control plane Kubernetes controllers, Knative or HPA, Gateway API KubeRay plus Ray scheduler and Serve controller BentoCloud or self-managed operator
Autoscaling signal Concurrency, RPS, CPU, or LLM metrics via KEDA/HPA Ongoing requests per replica Concurrent requests over declared concurrency
Scale to zero Yes in Knative mode, or with KEDA Possible with min_replicas: 0, higher tail latency Yes with external queue on BentoCloud
Multi-model pipelines Transformer plus predictor, or InferenceGraph Strongest: handles between deployments Service composition in Python
LLM routing and disaggregation Documented in v0.17 with llm-d style components Documented in Ray Serve LLM, engine-dependent Not documented in sources reviewed
GPU sharing Whatever Kubernetes and the device plugin provide Fractional num_gpus, no memory isolation Resource declaration, scheduler dependent
Canary and rollback canaryTrafficPercent, Knative revisions Standby cluster upgrades via KubeRay Redeploy earlier Bento tag
Ops burden Highest if you self-run Knative and gateways High: two schedulers to understand Lowest on BentoCloud, higher self-hosted
Developer onboarding Needs Kubernetes literacy Needs Ray literacy Fastest for Python teams
Best fit Platform teams serving many tenants Ray-centric orgs, complex pipelines Small teams optimizing velocity

Read the matrix by rows that matter to you, not by column totals. A team that serves forty independent models for twelve internal tenants should weight “control plane” and “canary” heavily and will likely land on KServe. A team whose product is a multi-stage retrieval pipeline should weight “multi-model pipelines” and will find Ray Serve natural. A four-person team with two production models should weight “onboarding” and “ops burden.”

Decision flow for choosing KServe, Ray Serve or BentoML for model serving on Kubernetes

Figure 4: A decision flow driven by workload shape and team ownership rather than feature lists.

The flow in Figure 4 encodes two questions that outperform feature comparison: who owns Kubernetes in your organization, and whether Ray is already in the building. Both are organizational facts, and organizational facts dominate in practice.

Trade-offs, Gotchas, and What Goes Wrong

Cold starts versus scale-to-zero. All three can scale to zero, and all three make it dangerous for large models. KServe’s documentation cites roughly 10 seconds of cold start in its examples, which is plausible for a small container with a cached image. A multi-gigabyte weight download plus CUDA initialization plus engine warm-up for a large LLM can take minutes. Requests that arrive in that window either queue (BentoML’s external queue explicitly pauses the timeout timer until a replica is ready) or time out. Mitigate with a nonzero minimum for anything user-facing, with node-local model caches (KServe’s localmodel component exists for this purpose), and with pre-pulled images.

Autoscaling on the wrong signal. Concurrency-based scaling assumes comparable request cost. For LLMs it is a poor proxy: ten concurrent requests with long contexts can saturate a GPU that would happily handle fifty short ones. This is why KServe’s LLM autoscaling watches KV cache utilization and queue depth, and why a plain BentoML concurrency formula or a default Ray target_ongoing_requests needs LLM-specific tuning. If your scaler does not see KV cache pressure, it will scale late and then overshoot.

Two schedulers, two failure domains (Ray). When a Ray Serve replica is pending, the cause could be Serve (no capacity under its limits), Ray (no free resource in the cluster), KubeRay (worker group at max), or Kubernetes (no node or quota). Teams without a runbook for those four layers lose hours. Also remember the upgrade behavior: a standby-cluster rollout temporarily doubles the footprint, and a cluster with no spare GPUs cannot complete it.

Fractional GPUs without isolation. Ray’s num_gpus: 0.5 is bookkeeping. If two replicas each load a model and each grows its activation memory under load, one will hit an out-of-memory error and the failure will look random. Use MIG for hard isolation where the hardware supports it, or leave headroom and cap batch sizes.

Operational weight (KServe). Serverless mode brings Knative and a mesh or ingress layer; LLM mode brings Gateway API, the Inference Extension and optionally Envoy AI Gateway and RDMA-capable networking for disaggregated serving. The v0.17 Helm restructuring into ten charts and the explicit no-simple-upgrade warning are a reminder that this is infrastructure you will own. The upside is that none of it is proprietary.

Managed-versus-self-hosted feature gaps (BentoML). The features in BentoML’s autoscaling documentation are described in the context of BentoCloud. If you plan to self-host on your own cluster for compliance reasons, validate that scale-to-zero, the external queue and stabilization policies are available to you before you design around them. I could not confirm the self-hosted parity from the sources reviewed, so I flag it as unverified.

Batching conflicts. Adaptive batching (BentoML), dynamic batching (Ray Serve) and the runtime’s own batching (Triton, vLLM) can stack. Two layers of batching add latency without adding throughput, and each waits for the other. Pick one owner for batching per model and disable the rest.

Protocol lock-in. If many clients call your models, the wire contract matters more than the platform. KServe’s Open Inference Protocol and OpenAI-compatible LLM endpoints are the most portable choices; Ray Serve and BentoML let you define arbitrary HTTP APIs, which is flexible and also the easiest way to create an unversioned contract that breaks when you migrate.

Practical Recommendations

Start with the organizational question, then the workload question. Platform-team-owned Kubernetes with many tenants and strict governance points to KServe. Existing Ray investment and graph-shaped inference pipelines point to Ray Serve. A small team with a Python-first culture and a preference for managed operations points to BentoML.

For LLM-heavy fleets specifically, I would favor KServe’s LLMInferenceService if you can staff the Gateway and networking layers, because the routing and autoscaling components are being standardized in the open and align with the broader llm-d work. I would favor Ray Serve LLM if you already run Ray or need unusual routing and composition logic. I would use BentoML for LLMs mainly when the deployment is modest and developer speed matters more than fleet-level cache-aware routing.

Mixed estates are legitimate. Many organizations run KServe for the long tail of standard models and Ray Serve for one or two complex pipelines. The cost is two runbooks; the benefit is not forcing a pipeline into a CRD or a tenant model into a Ray cluster. If you go this route, standardize at the edge: one gateway, one authentication scheme, one metrics namespace.

Whichever you choose, run a one-week proof of concept with your real model, your real traffic shape and your real cold start, and record four numbers: p99 latency at steady load, time from traffic spike to new replica serving, rollout capacity overhead, and the number of distinct components an on-call engineer must understand. Those four numbers settle more debates than any comparison table. For event-driven scaling approaches that work across all three, see our guide to KEDA event-driven autoscaling on Kubernetes.

Checklist before you decide

  • Name the owner of the request queue and the concurrency limit for each model.
  • Measure real cold start, including weight load and warm-up, before enabling scale-to-zero.
  • Choose one layer to own batching and disable the others.
  • Decide your canary mechanism and verify rollback takes one change, not a redeploy.
  • Confirm GPU isolation needs: MIG, time-slicing or fractional bookkeeping.
  • For LLMs, confirm the autoscaler sees KV cache pressure or queue depth, not only request counts.
  • Pin versions and read the upgrade notes; KServe 0.17 needed a documented migration, not helm upgrade.
  • Standardize the client-facing protocol so the platform can change later.

Frequently Asked Questions

Is KServe or Ray Serve better for LLM serving?

Neither is universally better. KServe’s LLMInferenceService became production-ready in version 0.17 and bundles cache-aware routing, prefill/decode disaggregation and token-based rate limiting as Kubernetes resources, which suits platform teams. Ray Serve LLM offers a Python programming model with vLLM and SGLang integration, prefix-aware routing and disaggregation patterns, which suits teams already on Ray or needing custom composition. Choose by who operates the cluster and how custom your pipeline is.

Can BentoML run on Kubernetes without BentoCloud?

Yes, you can containerize a Bento and run it on your own cluster, and BentoML has historically offered Kubernetes deployment paths. However, the autoscaling, scale-to-zero and external queue features described in BentoML’s documentation are presented in the context of BentoCloud. Before relying on them in a self-hosted setup, confirm with the current docs which features are available and which you must build yourself with tools such as HPA or KEDA.

Does KServe support scale to zero?

Yes. In Knative serverless mode, setting minReplicas: 0 allows pods to terminate when idle, and the Knative autoscaler uses a 60-second stable window with a 6-second panic window per KServe’s documentation. In RawDeployment mode HPA cannot scale to zero natively, so teams use KEDA. For large models, weigh minutes-long cold starts against the cost savings before enabling it on user-facing endpoints.

How does Ray Serve autoscaling work on Kubernetes?

Ray Serve scales replicas based on target_ongoing_requests, which defaults to 2 since Ray 2.32, bounded by min_replicas and max_replicas with an upscale_delay_s default of 30 seconds. On Kubernetes, KubeRay separately scales Ray worker pods so there is room for those replicas. That means two autoscalers act in series: tune and monitor both, or a burst will wait on the slower layer.

Which platform has the best canary deployment support?

KServe has the most explicit canary primitive: canaryTrafficPercent in Knative mode splits traffic between the latest and previous revisions. Ray Serve relies on KubeRay’s standby-cluster upgrade, which switches traffic after the new cluster is healthy but may temporarily double GPU use. BentoML versions artifacts, so rollback is redeploying an earlier tag, with traffic splitting depending on the deployment platform you use.

Should I use all three together?

Only if you have a reason. A mixed estate, for example KServe for the long tail of standard models and Ray Serve for one complex pipeline, is defensible when each owns a distinct workload class. Running all three for the same kind of workload multiplies runbooks, dashboards and upgrade work. Standardize the gateway, authentication and metrics at the edge so that clients and on-call engineers see one consistent surface.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *