llm-d on Kubernetes: Prefill/Decode Disaggregation for LLM Inference

llm-d on Kubernetes: Prefill/Decode Disaggregation for LLM Inference

llm-d on Kubernetes: Prefill/Decode Disaggregation for LLM Inference

Running a large language model behind a plain Kubernetes Service is the fastest way to discover that inference traffic is not web traffic. Two requests that look identical to a load balancer can differ by a factor of a hundred in GPU work, and a round-robin policy will happily stack a 30,000-token document summary on the same replica that is midway through forty chatty sessions. llm-d Kubernetes deployments attack that mismatch directly: the project, launched in May 2025 and accepted into the CNCF Sandbox on 24 March 2026, layers an inference-aware router, prefill/decode disaggregation and tiered KV-cache management on top of vLLM. It is now at v0.10.0, released on 29 September 2026.

This matters because GPU time, not model quality, has become the dominant line item in most production LLM budgets, and the gap between a naive deployment and a well-routed one is measured in multiples rather than percentages. In this post you will learn how the llm-d architecture works, what each component does, when splitting prefill from decode pays for itself (and when it does not), how to reason about KV-transfer bandwidth and cost per token, and what an installation path looks like.

What this covers: the problem llm-d solves, its reference architecture, a walk-through of routing and disaggregation, worked cost and bandwidth math, an install sketch, failure modes, and a decision checklist.

Context and Background

For most of 2023 and 2024, “serving an LLM on Kubernetes” meant a Deployment of vLLM, Hugging Face TGI or NVIDIA Triton pods behind a Service or an Ingress. That works surprisingly well for small models and uniform traffic, and it is where most teams should start. Our comparison of vLLM, SGLang and TensorRT-LLM in 2026 covers the engine layer, and the benchmark follow-up shows how much the choice of engine moves single-replica numbers.

What single-engine benchmarks cannot show is fleet behaviour. Once you run tens of replicas, three problems appear. First, load balancers see requests as opaque HTTP calls, so they cannot know which replica already holds the KV cache for a shared system prompt. Second, prefill (processing the prompt) and decode (generating tokens one at a time) have opposite hardware profiles, yet they share one GPU and interfere with each other. Third, the KV cache, which grows linearly with context length, becomes the binding constraint long before compute does.

The industry response has been to treat inference as a distributed-systems problem. Research systems such as DistServe (OSDI 2024) showed that assigning prefill and decoding to different GPUs removes interference between the phases; the authors report serving up to 7.4 times more requests, or meeting 12.6 times tighter latency objectives, than existing systems on their workloads (see the DistServe paper). Those numbers come from a specific research setup and should be read as an upper-bound illustration, not a forecast for your cluster.

llm-d is the Kubernetes-native packaging of these ideas. It was founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA, with AMD, Cisco, Hugging Face, Intel, Lambda and Mistral AI among later supporters, according to the CNCF announcement. It is Apache 2.0 licensed and deliberately does not replace vLLM: vLLM (or SGLang) remains the model server, while llm-d contributes the router, the orchestration recipes and the operational glue. That framing is the key to the “llm-d vs vLLM” question, which we return to later: they are complements, not competitors.

The standard that makes the routing layer portable is the Kubernetes Gateway API Inference Extension (GAIE). It defines an InferencePool resource, which groups model-server endpoints, and an Endpoint Picker protocol built on Envoy’s external processing (ext-proc) filter, letting a gateway ask a specialised component which replica should serve each request. For the broader gateway picture see our comparison of Gateway API versus Ingress.

The llm-d Reference Architecture

llm-d splits an LLM serving cluster into a router that understands inference state and pools of model servers that can be specialised by phase. The router picks endpoints using prefix-cache and load signals, and can orchestrate a two-step prefill-then-decode path, while KV blocks move between workers over RDMA-capable transports such as NIXL.

llm-d Kubernetes reference architecture with router, prefill pool and decode pool

Figure 1: llm-d Kubernetes reference architecture. Clients reach a gateway, the llm-d Router selects endpoints, and separate prefill and decode pools exchange KV cache blocks.

The diagram shows the request path. A client sends an OpenAI-compatible request to a gateway or standalone Envoy proxy. The proxy consults the Endpoint Picker (EPP), which is the brain of the llm-d Router. The EPP scores candidate endpoints, and the proxy forwards the request to the winner. In a disaggregated deployment, the router selects a prefill worker and a decode worker, and a routing sidecar on the decode pod coordinates the handoff of KV cache blocks from the prefill worker.

The three layers: router, pool, model server

The llm-d documentation describes three primary elements. The Router combines a proxy with the Endpoint Picker. The InferencePool is the Kubernetes object that groups the model-server pods the router may choose from. The Model Server is the inference engine itself, typically vLLM, with SGLang supported on NVIDIA GPUs for the disaggregation guide.

Separating these layers is a sound design choice for a practical reason: each changes at a different speed. Engines ship a release every few weeks and change flags constantly. The routing policy changes when your traffic changes. The pool definition changes when you add hardware. Because the EPP talks to the pool through a metrics-and-capabilities interface (prefix-cache state, queue depth, LoRA adapters loaded), you can upgrade vLLM without re-plumbing the gateway.

Since v0.7 (May 2026) the default deployment has shifted towards a standalone proxy mode, where the router runs with an Envoy sidecar and does not require a full Kubernetes Gateway. Gateway mode remains available with providers such as GKE Gateway, Istio and agentgateway. If you already run a service mesh, our piece on Istio Ambient, multicluster and the Gateway API Inference Extension explains how the same extension plugs into a mesh data plane.

Why the router is the highest-leverage piece

If you adopt only one part of llm-d, adopt the router. The project’s own experiments on two NVIDIA 8xH100 nodes report roughly three times lower mean time-to-first-token (TTFT) with prefix-aware routing at four queries per second on two replicas, about 50 percent more sustained queries per second under a P95 TTFT limit of two seconds on four replicas, and about twice the baseline throughput for Llama 3.1 70B in FP16 on four replicas. These are vendor-published results on synthetic prefix-heavy workloads; your gain depends entirely on how much prefix overlap your traffic has.

The v0.5 release blog reports a Qwen3-32B run on eight vLLM pods across 16 H100 GPUs (tensor parallelism of two), with up to 109 percent higher throughput and 99 percent lower TTFT than a baseline Kubernetes service. The CNCF post describes the same class of test as staying near zero latency while scaling towards roughly 120,000 tokens per second, where the baseline degrades rapidly under load. Treat “99 percent lower TTFT” as the behaviour of a saturated baseline queue, not as a steady-state latency claim.

The mechanism behind these gains is simple. A round-robin balancer creates a feedback loop: a busy replica produces slower tokens, its connections stay open longer, and connection-count-based balancers may then send it even more work. A router that reads queue depth and KV-cache utilisation directly breaks that loop, and one that also knows which replica already holds a prompt prefix avoids redundant prefill entirely.

What disaggregation adds

Disaggregation moves prefill and decode to different pods. Prefill is compute-bound: it processes every prompt token in parallel, so it saturates tensor cores. Decode is memory-bandwidth-bound: each step reads all model weights and the accumulated KV cache to produce a single token per sequence. When both run on one GPU, a large prefill stalls in-flight decodes, which shows up as spikes in inter-token latency (ITL).

vLLM’s documentation states the trade-off candidly: disaggregated prefill lets you tune TTFT and ITL separately and controls tail latency, but it does not by itself improve throughput. The llm-d guide makes the same point more carefully: for a fixed ITL target, disaggregation can raise throughput per GPU because you can specialise each pool and use wide parallelism on decode. It is explicitly not a target for every workload, and the guide recommends considering it for medium-to-large models, long inputs such as 10,000 input tokens to 1,000 output tokens, and sparse mixture-of-experts architectures.

Deeper Analysis: Routing, KV Transfer and the Cost Math

The interesting engineering in llm-d sits in three places: how the scheduler scores endpoints, how KV blocks physically move between pods, and how those two decisions translate into cost per token. This section walks through each with worked numbers. Where a figure is my own arithmetic rather than a published measurement, it is labelled illustrative.

How KV-cache-aware routing scores an endpoint

KV cache aware routing means the scheduler prefers the replica that already holds the KV blocks for the incoming prompt prefix, so prefill work for that prefix is skipped. The EPP does this by combining several scorers in a weighted sum: a prefix-cache scorer, a load scorer (queue depth and KV utilisation), and optionally a predicted-latency scorer. The llm-d docs describe an XGBoost-based predictor that estimates latency per endpoint from live features, which makes the policy adaptive rather than hand-tuned.

KV cache aware routing decision flow in llm-d Kubernetes scheduler

Figure 2: KV cache aware routing flow. The EPP filters unhealthy endpoints, scores the rest on prefix hit and load, and picks the highest total score.

The flow in Figure 2 has a subtle property worth stating: prefix affinity and load balancing pull in opposite directions. If every request shares a 2,000-token system prompt, pure prefix affinity sends all traffic to one replica, which then saturates. Weighted scoring lets a heavily loaded replica lose despite a perfect cache hit, and the cache warms elsewhere as a side effect. The weights are the tuning surface, and the right values depend on how expensive your prefill is relative to queueing delay.

Prefix awareness depends on the router knowing what each replica has cached. llm-d supports a precise mode, where an indexer tracks KV block events emitted by vLLM, and lighter approximate modes that estimate cache contents from routing history. Precise tracking costs some control-plane traffic but avoids the drift that makes approximate schemes misroute after evictions. As of v0.10.0, the release notes record that the KV-cache indexing code has moved from the separate llm-d-kv-cache repository into llm-d-router, so older tutorials that reference the previous repository layout are out of date.

Sizing the KV cache: the number that drives everything

KV cache size per token is 2 x layers x KV heads x head dimension x bytes per element. For a Llama-3.1-70B-class model (80 layers, 8 KV heads under grouped-query attention, head dimension 128) in 16-bit precision, that is 2 x 80 x 8 x 128 x 2 = 327,680 bytes, or 320 KiB per token. This uses the model’s published architecture; the arithmetic is mine.

A 10,000-token prompt therefore carries roughly 3.3 GB of KV cache (illustrative, 16-bit KV). On an 80 GB GPU that also holds a share of 140 GB of FP16 weights, the cache budget is measured in tens of gigabytes, so only a handful of long contexts fit per replica. This is why context length, not compute, is usually the first wall you hit, and why tiered offloading is a first-class llm-d feature.

The v0.5 blog reports a hierarchical-offload test on Llama-3.1-70B with four H100 GPUs and IBM Storage Scale, using 16K-token requests with cached prompts: about 185,000 tokens per second at 250 concurrent users, described as a 13.9 times improvement over GPU-only serving at saturation. The workload is cache-hit-heavy by construction, so read it as the ceiling for repeated-context workloads such as agents and long multi-turn chats, not as a general result.

What the KV transfer costs

Disaggregation introduces a transfer that colocated serving never pays. Using the 3.3 GB example, the ideal wire time depends on the fabric (my arithmetic, ignoring protocol overhead and assuming the link is dedicated):

Fabric (per node) Usable bandwidth Transfer time for 3.3 GB
400 Gb/s RDMA NIC (InfiniBand or RoCE) about 50 GB/s about 66 ms
100 Gb/s RDMA NIC about 12.5 GB/s about 260 ms
25 Gb/s TCP about 3 GB/s about 1.1 s

Table: illustrative KV transfer time for a 10,000-token Llama-3.1-70B-class context, 16-bit KV. Real throughput is lower once congestion, small-block overhead and shared links are included.

Compare that with the compute you saved. Prefill for 10,000 tokens on a 70B dense model is about 2 x 70e9 x 10,000 = 1.4e15 floating-point operations. At an assumed effective 400 TFLOP/s per H100-class GPU (illustrative), that is around 3.5 seconds on one GPU, or roughly 0.9 seconds spread over four GPUs at tensor parallelism of four. A 66 ms transfer is a rounding error against that; a one-second TCP transfer erases the benefit for this prompt size. The lesson: disaggregation is a network-fabric decision as much as a scheduling one.

NIXL (NVIDIA Inference Xfer Library) is the default transport in llm-d’s guides, riding on UCX to use RDMA or fall back to TCP. Mooncake’s transfer engine is an alternative connector with a stricter rule: prefill and decode must use the same tensor-parallel size. The llm-d docs also quantify one operational cost: establishing a NIXL connection between a decode and a prefill instance takes roughly five seconds, paid once per pair, which llm-d handles with a “dynamic lazy” strategy instead of a central bootstrap server.

The request path in a disaggregated deployment

Sequence of a disaggregated request through llm-d prefill and decode workers

Figure 3: sequence of one disaggregated request. The router picks both workers, prefill produces KV blocks, and the decode worker pulls them before generating tokens.

Figure 3 shows why a sidecar exists. The decode pod runs a routing proxy that receives the request, calls the chosen prefill worker with a request that produces KV blocks only, then instructs the local vLLM to pull those blocks and stream tokens back. The client sees one OpenAI-compatible call. In vLLM terms the prefill worker acts as a kv_producer and the decode worker as a kv_consumer.

One design detail matters in practice: the decision to disaggregate is itself a pluggable step. Short prompts, or prompts whose prefix is already cached on a decode worker, are not worth splitting, because the transfer overhead can exceed the prefill cost. The llm-d Endpoint Picker configuration models this as a P/D decider plugin. The reference guide uses always-disagg-pd-decider, which splits every request because its benchmark workload has long inputs; the plugin interface exists so that you can replace it with a policy that keeps short or well-cached requests on a single worker. Check the current plugin list in the docs before relying on a specific alternative, since the router is evolving quickly.

A worked cost per token example

Cost per million output tokens equals hourly cluster cost divided by tokens per hour, times one million. Assume an illustrative on-demand price of $3.00 per H100-hour (real prices vary widely by provider and commitment). The guide’s reference topology of 8 prefill instances at TP=1 and 2 decode instances at TP=4 uses 16 GPUs, or $48 per hour.

Sustained output rate Tokens per hour Cost per million output tokens
2,000 tok/s 7.2 million about $6.67
4,000 tok/s 14.4 million about $3.33
6,000 tok/s 21.6 million about $2.22

Table: illustrative cost per million output tokens for a 16-GPU cluster at $3.00 per GPU-hour. These are not benchmarks of llm-d; they show the sensitivity of unit cost to utilisation.

The point of the table is the sensitivity. Moving from 2,000 to 6,000 tokens per second cuts unit cost by two thirds without touching the hourly bill. Routing improvements that raise sustained throughput, and disaggregation that lets you hit an ITL target with fewer GPUs, act on this denominator. If you are pricing inference internally, sustained utilisation under your latency SLO is the number to measure; peak tokens per second on a quiet cluster is not.

Reading the published wide expert-parallel result

llm-d’s “wide expert parallelism” path targets large mixture-of-experts models. The v0.5 blog reports a run with 16 prefill GPUs and 16 decode GPUs on NVIDIA B200 (EP=16, DP=16, TP=1) reaching about 50,000 output tokens per second in total, roughly 3,100 tokens per second per decode GPU. That figure is a headline for a heavily tuned recipe on top-end hardware and a matching network; it demonstrates that the architecture scales, not what a four-GPU team should expect. The blog also reports a UCCL host-driven transport that degraded 7.1 percent under network contention versus 17.1 percent for the comparison, a reminder that KV traffic competes with expert-parallel all-to-all traffic on the same fabric.

Walk-through: Installing llm-d and Reading the Config

The llm-d project calls its tested deployment recipes “well-lit paths”: an optimized baseline (inference scheduling), tiered prefix caching, prefill/decode disaggregation, and wide expert parallelism. Each is a guide in the repository with Helm values for the router and Kustomize overlays for the model server. The steps below follow the P/D disaggregation guide as it stands at v0.10.0; always check the guide on the tag you deploy, because flags and chart names change between releases.

Prerequisites

You need a Kubernetes cluster with GPU nodes and a GPU device plugin or DRA driver, kubectl, helm and git, plus a Hugging Face token stored as a secret. For real disaggregation you also want RDMA-capable networking exposed inside pods. The guide’s GKE recipe uses Dynamic Resource Allocation for GPUs and DRANET for RoCE; on CoreWeave or other clouds you would use the matching overlay. If host-level RDMA works but pods cannot see the devices, KV transfer falls back to slow paths or fails.

# 1. Get the guides at a pinned tag (v0.10.0 at the time of writing)
git clone https://github.com/llm-d/llm-d.git && cd llm-d
git checkout v0.10.0
export REPO_ROOT=$(realpath $(git rev-parse --show-toplevel))
source ${REPO_ROOT}/guides/env.sh

export GUIDE_NAME="pd-disaggregation"
export NAMESPACE="llm-d-pd-disaggregation"
export MODEL_NAME="openai/gpt-oss-120b"

# 2. Gateway API Inference Extension CRDs (InferencePool and friends)
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api-inference-extension/${GAIE_URL}/v1-manifests.yaml

# 3. Namespace and model-download secret
kubectl create namespace ${NAMESPACE}
kubectl create secret generic llm-d-hf-token \
  --from-literal="HF_TOKEN=${HF_TOKEN}" -n ${NAMESPACE}

# 4. Router in standalone mode (Envoy sidecar, no Gateway object needed)
helm install ${GUIDE_NAME} ${ROUTER_STANDALONE_CHART} \
  -f ${REPO_ROOT}/guides/recipes/router/base.values.yaml \
  -f ${REPO_ROOT}/guides/${GUIDE_NAME}/router/${GUIDE_NAME}.values.yaml \
  -n ${NAMESPACE} --version ${ROUTER_CHART_VERSION}

# 5. Model servers: 8 prefill pods at TP=1 and 2 decode pods at TP=4
export INFRA_PROVIDER=base   # or coreweave, gke/base, aws ...
kubectl apply -n ${NAMESPACE} \
  -k ${REPO_ROOT}/guides/${GUIDE_NAME}/modelserver/gpu/vllm/${INFRA_PROVIDER}

The 8-prefill, 2-decode split is instructive. Prefill pods run at TP=1 because prefill parallelises well across replicas and does not need wide tensor parallelism, while decode pods use TP=4 so that each decode instance has more aggregate memory bandwidth and KV space. This heterogeneous-parallelism pattern is a core reason disaggregation can beat colocated serving.

Reading the router configuration

The scheduling behaviour lives in the Endpoint Picker’s plugin configuration. The excerpt below is abbreviated from the guide’s values file; comments are mine.

apiVersion: llm-d.ai/v1alpha1
kind: EndpointPickerConfig
plugins:
- type: always-disagg-pd-decider        # split every request
- type: disagg-profile-handler
  parameters:
    deciders:
      prefill: always-disagg-pd-decider
- type: prefill-filter                  # only prefill pods
- type: decode-filter                   # only decode pods
- type: approx-prefix-cache-producer    # estimate cache contents
- type: inflight-load-producer          # track in-flight tokens
- type: prefix-cache-affinity-filter    # keep prefix groups warm
- type: token-load-scorer               # queued prefill tokens
- type: active-request-scorer           # concurrent requests
- type: max-score-picker
schedulingProfiles:
- name: prefill
  plugins: [prefill-filter, prefix-cache-affinity-filter,
            token-load-scorer, max-score-picker]
- name: decode
  plugins: [decode-filter, active-request-scorer, max-score-picker]

Two scheduling profiles run per request, one for each phase, and each uses a different load signal. Prefill selection is driven by queued prompt tokens, because prefill cost is proportional to prompt length. Decode selection is driven by the number of concurrent requests, because decode throughput is bound by how many sequences share each memory-bound step. Picking a single “least loaded” metric for both phases is a common mistake that this configuration avoids.

Note also the peakPrefillThroughput parameter in the guide’s values, which the affinity filter uses as a gate. The comments in the file state it was measured with a calibration script through the full P/D path on a specific reference fleet and must be re-measured for other hardware and models. Copying the number blindly to different GPUs is a silent way to misconfigure the router.

llm-d Kubernetes deployment topology showing router, InferencePool, prefill and decode pods and KV tiers

Figure 4: deployment topology. The router fronts an InferencePool of prefill and decode pods; a KV tier of CPU memory and storage sits behind the pods, and an autoscaler watches queue and latency signals.

Autoscaling and the rest of the platform

Static pool sizes are the weakest point of any disaggregated design, because the right prefill-to-decode ratio (the “xPyD” ratio) shifts with the mix of input and output lengths. llm-d supports HPA and KEDA-based autoscaling, and a variant autoscaler that has, as of v0.10.0, been renamed from the workload-variant-autoscaler repository to llm-d-autoscaling; the older WVA guides are marked deprecated. The v0.5 blog also describes scale-to-zero with a cold-start activator for intermittent workloads.

Multi-node replicas and expert parallelism use the LeaderWorkerSet (LWS) API. The P/D guide also offers a DisaggregatedSet variant, which manages the whole prefill/decode topology as a single LWS resource with coordinated rollouts, which addresses a genuine operational problem: upgrading a prefill pool and a decode pool independently can leave the two on incompatible versions mid-rollout, and KV blocks are not portable across mismatched model or engine configurations.

Other features are worth knowing but should be adopted separately. Flow control provides queueing and priority admission at the router. Batch processing exposes OpenAI-compatible batch APIs through a batch gateway. Multimodal serving, batch and flow control were marked as graduated to production status in the v0.8.0 notes. Intel XPU, AMD, TPU and MetaX paths exist, but the guide is explicit that non-default hardware variants often use reduced configurations maintained by their vendors and are not guaranteed production sizing.

Trade-offs, Gotchas, and What Goes Wrong

Disaggregation adds moving parts, and every moving part fails in its own way. The most useful way to plan is to list the failure modes before the first deployment.

KV transfer failures and silent recomputation. If a transfer fails and the decode worker quietly recomputes the prompt, requests still return HTTP 200 and you have paid for disaggregation without receiving it. The guide’s MetaX notes recommend kv_load_failure_policy=fail so a failed pull errors out instead, and suggest verifying decode logs for an external prefix-cache hit. Do the same in your own validation: assert that transfers happen, do not infer it from success codes.

Tensor-parallel asymmetry. The NIXL connector supports only the fan-out direction: decode TP must be greater than or equal to prefill TP, for most model architectures. Prefill TP=8 with decode TP=4 is explicitly unsupported unless the model is MLA or Mamba-based. That constraint shapes topology choices: you cannot freely give prefill the wide parallelism and decode the narrow one.

Stale connections and stranded KV. The docs list known NIXL issues, including stale agent caching after a prefill pod restarts. On the SGLang path with NIXL, a request cancelled before the decode side starts its transfer can strand KV cache on the prefill worker until the pod restarts, because there is no prefill-side reclaim timeout. Mooncake’s connector documents a configurable abort timeout (default 480 seconds) after which a prefiller releases KV if the decoder never acknowledges. Cancelled requests are therefore a real operational scenario: load-test them.

Network as a shared resource. KV transfer, expert-parallel all-to-all and storage-backed offload can all contend for the same NICs. A 3 GB transfer that takes 66 ms on an idle 400 Gb/s link takes far longer when a wide-EP job shares it. Measure P99 transfer time, not the average.

Small models and short prompts. For a 7B to 8B model with 200-token prompts, the transfer, the extra hop and the extra pods dominate any interference savings. The llm-d guide says as much. Prefix-aware routing alone, which needs no special network, often captures most of the benefit for such workloads.

Experimental surfaces and churn. vLLM labels its disaggregated prefill feature experimental. The llm-d project itself is a young CNCF Sandbox project on a roughly monthly release cadence (v0.6 in April, v0.7 in May, v0.8 in June, v0.9 in August and v0.10 in late September 2026), with repository migrations and deprecations in the latest release. Pin versions, read release notes, and budget engineering time for upgrades.

Observability gaps. Aggregate GPU utilisation hides the phase imbalance. Track TTFT, time per output token (TPOT), queue depth per pool, KV utilisation, cache-hit rate, and transfer latency separately for prefill and decode. If prefill pods are saturated while decode pods idle, your ratio is wrong, and no dashboard averaging both will show it.

Benchmark provenance. Nearly all published llm-d numbers come from the project and its sponsors, run on specific models, hardware and synthetic workloads. That does not make them wrong, but it means you should replay your own traffic traces before committing capital.

Practical Recommendations

Adopt llm-d in stages, and let measurement, not architecture enthusiasm, decide how far you go. The staging below matches how the project’s own guides are layered: each step is independently useful and each adds operational surface.

Start with the router. Deploy the optimized-baseline guide in front of your existing vLLM pool, keep the model servers unchanged, and compare TTFT and throughput under your own traffic replay. If your workload has shared system prompts, RAG templates, agent scaffolds or multi-turn sessions, prefix-aware routing tends to be the cheapest win, because it needs no RDMA fabric and no topology changes. Keep the version of the Gateway API Inference Extension CRDs aligned with the version the guide expects.

Add tiered KV offload next if your contexts are long and your cache hit rate is high but GPU memory keeps evicting. Offload trades latency on a cache hit for capacity, so measure the hit path, not only the miss path.

Only then consider disaggregation, and only if all of the following hold: the model is medium-to-large or sparse MoE, inputs are long relative to outputs, you have a genuine ITL or TTFT objective that colocated serving misses, and you have RDMA-capable networking visible inside pods. If any one is false, the likely outcome is more complexity for little gain.

Checklist before you go to production:

  • Pin llm-d, vLLM and GAIE versions; read each release’s deprecations (v0.10.0 deprecates the llm-d-cuda image in favour of upstream vllm/vllm-openai).
  • Validate RDMA from inside a pod before deploying, not only on the host.
  • Keep decode TP greater than or equal to prefill TP when using NIXL.
  • Re-measure peakPrefillThroughput and similar calibration values on your hardware.
  • Assert real KV transfers in tests (fail policy on, logs checked), including cancelled and restarted-pod scenarios.
  • Dashboard TTFT, TPOT, per-pool queue depth, KV utilisation, cache hit rate and transfer latency.
  • Compute cost per million output tokens at sustained load under your SLO, and compare against a tuned colocated baseline, not a naive one.
  • Keep a rollback path to a plain vLLM Deployment.

Frequently Asked Questions

What is llm-d and is it a replacement for vLLM?

llm-d is a Kubernetes-native distributed inference stack, an Apache 2.0 CNCF Sandbox project. It is not a replacement for vLLM. vLLM (or SGLang) remains the model server that runs the GPU kernels, while llm-d adds an inference-aware router, prefill/decode disaggregation recipes, KV-cache management, and autoscaling. You can think of vLLM as the engine and llm-d as the orchestration and traffic layer wrapped around a fleet of those engines.

llm-d vs vLLM: when is plain vLLM enough?

Plain vLLM is enough for a single model on a handful of replicas with uniform, short requests, or for prototypes. It also suffices when prefix reuse is low and latency objectives are loose. llm-d begins to pay off when you run many replicas, share prompts across requests, serve long contexts, or need separate TTFT and ITL targets. Because llm-d builds on vLLM, adopting the router first is a low-risk migration path.

Does prefill/decode disaggregation always improve throughput?

No. vLLM’s documentation states that disaggregated prefill does not by itself improve throughput; it improves control over TTFT and inter-token latency and reduces tail spikes. llm-d’s guide says that for a fixed ITL target it can raise throughput per GPU through specialisation and wide parallelism, but recommends it for larger models, long inputs and MoE architectures. For small models and short prompts, transfer overhead usually outweighs the gain.

What network do I need for KV cache transfer?

Prefer RDMA: InfiniBand or RoCE at 100 to 400 Gb/s per node, exposed inside pods through a device plugin or DRA. NIXL over UCX can fall back to TCP, which the guides accept for functional testing but not for performance. As a rough calculation, a 10,000-token context for a 70B-class model has about 3.3 GB of 16-bit KV cache, which takes about 66 ms on an ideal 400 Gb/s link and over a second on 25 Gb/s TCP.

Is llm-d production-ready, and which hardware does it support?

The project reports production-status graduation for several features, including flow control, batch processing and multimodal serving as of v0.8.0, and v0.10.0 is themed around operational hardening. It is still a CNCF Sandbox project with monthly releases and active deprecations, so treat it as production-capable with disciplined version pinning. Documented hardware includes NVIDIA GPUs, AMD, Google TPUs and Intel XPU, though vendor-maintained variants may use reduced configurations.

How does llm-d relate to the Gateway API Inference Extension?

The Gateway API Inference Extension (GAIE) is the Kubernetes SIG project that defines the InferencePool resource and the Endpoint Picker protocol over Envoy ext-proc. llm-d builds on it: its Router is an Endpoint Picker implementation with richer scorers for prefix cache, load and predicted latency, plus disaggregation logic. GAIE ships only a lightweight reference picker for conformance, so llm-d supplies the production-grade scheduling policy on top.

Further Reading

Internal:

External:

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *