Gateway API Inference Extension with Istio Ambient Multicluster: A Reference Architecture for LLM Serving
A round-robin load balancer is quietly one of the most expensive components in a GPU fleet. It treats every pod as interchangeable, but two vLLM replicas serving the same model can differ by an order of magnitude in how long a new request will wait: one has a warm prefix cache and an empty queue, the other is at 95% KV-cache utilization with dozens of requests pending. The Gateway API Inference Extension (GIE) exists to close that gap by letting the gateway ask the model servers what they are doing before it picks one.
This matters now because the pieces finally line up. GIE’s InferencePool API is stable at v1, Istio ships an implementation of it, and Istio’s ambient multicluster mode reached beta in the 1.29 release (February 2026). Together they promise sidecarless mTLS between clusters and inference-aware routing inside each. Together they also come with sharp edges that the announcements skip.
You will leave with a working mental model of the request path, a reference architecture for two-cluster LLM serving, the exact status of each component, and the failure modes to design for.
What this covers: the problem with cache-blind balancing, how the endpoint picker works, how Istio wires it in, what ambient multicluster adds and where it stops, a worked capacity example, failure modes, and a rollout checklist.
Context and Background
Kubernetes Services were designed for stateless, short, roughly uniform requests. A Service spreads connections across endpoints using kube-proxy rules or an Envoy load-balancing policy such as round robin or least request. Least-request is a decent proxy for load when requests are similar in cost. LLM requests are neither short nor uniform. A prompt may be 200 tokens or 120,000, the response streams for seconds to minutes, and the cost of a request depends on hidden server state.
That hidden state is the KV cache. During prefill, a transformer computes key and value tensors for every prompt token and keeps them in GPU memory so decoding does not recompute them. Servers such as vLLM reuse these blocks across requests that share a prefix (a system prompt, a retrieved document, a conversation history). A request landing on a replica that already holds its prefix skips most of the prefill work; the same request landing elsewhere pays the full cost. Routing that ignores this is leaving both latency and GPU hours on the table. If you want the serving-engine side of the story, our comparison of vLLM, SGLang and TensorRT-LLM covers how each engine manages its cache.
The evidence that cache-aware routing matters is public. The llm-d project, which builds an inference scheduler on top of GIE, published a benchmark in which precise prefix-cache-aware scheduling cut P90 time-to-first-token from about 92.6 seconds under random scheduling to about 0.54 seconds, on a Qwen 32B-class model across eight vLLM pods with 16 H100 GPUs and a deliberately cache-friendly B2B workload with 6,000-token shared contexts (llm-d blog). Treat that as an upper bound: it is a vendor-run benchmark on a workload designed to reward affinity. Your traffic will show smaller gains, but the direction is reliable.
On the mesh side, Istio has spent two years moving from sidecars to ambient mode, where a per-node ztunnel handles L4 mTLS and optional waypoint proxies handle L7 policy. Ambient multicluster extends this across clusters. The Istio project announced in March 2026 that both ambient multicluster and its Gateway API Inference Extension support were at beta (CNCF announcement). Istio 1.31.0 followed on August 31, 2026, with multicluster leak and credential-rotation fixes that show the feature is being exercised in real environments.
For the ingress and routing primitives underneath, see our earlier pieces on Kubernetes Gateway API versus Ingress and on Gateway API 1.6 TCPRoute and UDPRoute. If you are still choosing a mesh data plane, the Istio ambient versus Linkerd decision record covers that question.
A note on what this post corrects
Two claims circulate that need care. First, “ambient multicluster is production ready”: Istio’s announcement blog says the beta is not ready for production use, while the install documentation describes multi-network as beta with caveats. We treat it as a pilot-grade feature. Second, “Istio waypoints route by KV cache”: as of the sources we could verify, Istio’s inference support runs at the ingress Gateway, and waypoint integration was listed as future work. We flag both again where they matter.
The Reference Architecture: Inference Gateway plus Ambient Mesh
In this architecture, the Gateway API Inference Extension turns an ordinary Kubernetes Gateway into an inference gateway. An HTTPRoute points at an InferencePool instead of a Service, and an endpoint picker (EPP) chooses the pod for each request using live model-server metrics such as queue depth, KV-cache utilization, prefix-cache hits and LoRA adapter presence.

Figure 1: Reference architecture. The gateway consults the endpoint picker over ext_proc, the picker scores pods using scraped model-server metrics, and traffic then travels to the chosen pod over the mesh.
Figure 1 shows five parts. The client sends an OpenAI-style request to a Gateway. An HTTPRoute matches it and names an InferencePool as the backend. The gateway’s Envoy proxy calls the endpoint picker through Envoy’s External Processing (ext_proc) filter. The picker returns a chosen endpoint, and Envoy forwards the request there. Meanwhile the picker keeps a rolling view of every pod in the pool by scraping their metrics endpoints.
The InferencePool: a backend that knows it is an inference backend
An InferencePool is a Kubernetes custom resource in the inference.networking.k8s.io group. It selects model-server pods by label, names the ports they serve on, and points at the endpoint picker Service. The documented v1 shape looks like this:
apiVersion: inference.networking.k8s.io/v1
kind: InferencePool
metadata:
name: vllm-qwen3-32b
spec:
selector:
matchLabels:
app: vllm-qwen3-32b
targetPorts:
- number: 8000
endpointPickerRef:
name: vllm-qwen3-32b-epp
port:
number: 9002
failureMode: FailOpen
The pool is deliberately thin. It is a selector plus a pointer, not a scheduler. Everything smart lives in the picker, which is why the project can keep the API stable while scheduling algorithms churn. The API graduated to v1 with the v1.0.0 release, and one endpoint picker serves exactly one pool, although an HTTPRoute may reference several pools (InferencePool docs).
One consequence deserves emphasis: because the gateway knows the pool’s members, it can bypass the Service abstraction entirely. The Service VIP and kube-proxy are out of the data path for inference traffic, and the picker’s answer replaces the load-balancing decision.
The endpoint picker protocol
The EPP contract is small. The gateway streams request headers and body to the picker over gRPC ext_proc. The picker replies with the chosen address in the x-gateway-destination-endpoint header and in dynamic metadata. The gateway is required to track the pool’s endpoints, honour the returned choice, and report status per parent Gateway. If the picker cannot choose, the protocol defines outcomes: a 503 when no endpoints are ready and a 429 when the request should be dropped (implementer guide).
Why does the picker receive the body? Because prefix-aware scoring needs the prompt, and model-aware routing needs the model field in the JSON. A body-based routing step extracts that field so the route can match on model name or LoRA adapter, which is how one gateway hostname fronts many models.
What the picker actually scores
The GIE project used to bundle a full-featured picker. As of the v1.6.0 release, the project moved the full Endpoint Picker, Body-Based Routing and Latency Predictor components to the llm-d repository, and kept a minimal reference picker (lwepp) for conformance testing. It also removed the alpha InferenceObjective, InferenceModelRewrite and EndpointPickerConfig types from the project, again pointing at llm-d (v1.6.0 release notes). The project documentation now tells production users to bring their own picker or adopt one such as llm-d’s router.
That is a significant shift from the way many 2025 blog posts describe GIE, and it is worth internalising: the standard is the API and the protocol, and the scoring brain is a pluggable, separately maintained component. Note that Istio’s own task page still lists InferenceObjective among core resources, which suggests documentation lag rather than agreement; verify against the release you install.
The llm-d scheduler documents a filter, score, pick pipeline. Filters narrow candidates (by label, by prefill or decode role, by SLO tier). Scorers each return a value between 0 and 1: queue-depth-scorer prefers shorter queues, kv-cache-utilization-scorer prefers pods with more free cache, prefix-cache-scorer prefers pods matching the longest cached prefix, and lora-affinity-scorer prefers pods that already have the requested adapter loaded. The final score is the weighted sum of scorer outputs, so a scorer with weight 2.0 and output 0.8 contributes 1.6 (llm-d scheduling docs). A picker then takes the maximum, or samples from the scores to avoid herding.
Where Istio fits
Istio has supported GIE since 1.27 and, per its documentation, requires two pilot environment flags (ENABLE_GATEWAY_API_INFERENCE_EXTENSION and SUPPORT_GATEWAY_API_INFERENCE_EXTENSION) plus the Gateway API and inference CRDs. Istio’s docs pin Gateway API CRDs at v1.6.0 and the inference CRDs at v1.0.1 in their example. A DestinationRule is required for TLS between the gateway and the picker, since the picker is a gRPC service that should not be reachable in plaintext. Istio’s inference blog lists waypoint integration, HPA integration with model metrics, multi-modal optimisation and heterogeneous accelerators as future directions (Istio blog).
So the honest position is: in Istio, the inference-aware hop is at the edge gateway. Inside the mesh, ambient handles identity and encryption of the gateway-to-pod hop, but the routing intelligence is the picker’s, not the mesh’s.
Deeper Analysis: From One Cluster to Two
The single-cluster design in Figure 1 is where you should start. Multicluster adds capacity, failure isolation and data-residency options, but it also adds a second control loop that the endpoint picker knows nothing about. This section walks the request path, then layers ambient multicluster on top, and then quantifies what routing is worth.
The request path, step by step

Figure 2: Request lifecycle. The picker sits on the critical path of every request, which is why its latency and its failure mode both matter.
Figure 2 follows a single chat-completion call. The gateway terminates TLS and matches the route. It then opens a bidirectional ext_proc stream to the picker and sends headers and body. The picker refreshes or reads its cached view of the pool, runs filters and scorers, and returns one address. The gateway forwards the request, and the model server streams tokens back.
Three details are easy to miss. First, the picker’s view of the pool is stale by design: it scrapes metrics on an interval of tens to hundreds of milliseconds, so two requests arriving within one scrape interval can both be sent to the same “best” pod. Pickers mitigate this with in-flight accounting and probabilistic selection, but herding under bursts is a real effect. Second, the body must be buffered or streamed to the picker, so very large prompts add copy cost at the gateway; the project’s own v1.6.1 and v1.6.2 fixes concerned FULL_DUPLEX_STREAMED body handling and header ordering in the ext_proc stream, a sign this path is still maturing. Third, the response path can also be processed, for example to capture token usage, and that too is an extra hop.
The metrics come from the model servers themselves. vLLM exposes Prometheus metrics for waiting requests and KV-cache usage; the exact metric names have changed across vLLM releases (for example, GKE’s documentation refers to vllm:kv_cache_usage_perc for the current name). Pin your picker version to your model-server version, because a renamed metric silently turns a scorer into a constant.
Scoring in practice: a worked example
Consider three pods with the following state, and a picker configured with queue-depth weight 2, KV-cache weight 1 and prefix-cache weight 3. These numbers are illustrative, not measured.
| Pod | Queue score | Free KV score | Prefix score | Weighted total |
|---|---|---|---|---|
| A | 0.9 | 0.7 | 0.0 | 0.9×2 + 0.7×1 + 0.0×3 = 2.5 |
| B | 0.4 | 0.5 | 0.9 | 0.4×2 + 0.5×1 + 0.9×3 = 4.0 |
| C | 0.2 | 0.2 | 1.0 | 0.2×2 + 0.2×1 + 1.0×3 = 3.6 |
Pod B wins even though pod A is the emptiest, because B holds most of the request’s prefix and has tolerable queue depth. Pod C has the best prefix match but is nearly saturated, so its score falls behind. That trade between cache affinity and load is the entire design problem. Weight prefix too heavily and a popular system prompt turns one pod into a hotspot; weight queue depth too heavily and you have rebuilt least-request. The llm-d scheduler addresses the first failure with a prefix-cache-affinity-filter that applies probabilistic stickiness and falls back to load balancing on time-to-first-token, and with a no-hit-lru-scorer that spreads cold requests across the pool.
What routing is worth: a back-of-envelope model
Let a request carry a 6,000-token prompt where 5,000 tokens are a shared prefix. Suppose (illustratively) prefill processes 8,000 tokens per second per replica at moderate load. A full prefill costs 0.75 seconds. A cache hit on the 5,000-token prefix leaves 1,000 tokens to prefill, or 0.125 seconds. The saving is 0.625 seconds of GPU prefill time per request, an 83% reduction in prefill work.
With N replicas and blind balancing, the chance that a given replica holds a given hot prefix is roughly the fraction of replicas that have served it recently. If the working set of distinct prefixes is larger than one replica’s cache, blind routing thrashes: each replica evicts prefixes another would have reused. Cache-aware routing partitions the prefix space across replicas, so the effective fleet cache approaches the sum of the parts rather than a single replica’s capacity. That is the mechanism behind the very large llm-d numbers, and also the reason your gains depend on prefix reuse. For chat with unique prompts and no shared system prompt, prefix scoring adds nothing and only queue depth and KV pressure help.
Layering ambient multicluster
Now add a second cluster. Ambient multicluster in Istio uses a multi-primary, multi-network design: each cluster runs its own control plane, each network has an east-west gateway, and cross-cluster traffic is tunnelled over HBONE, Istio’s HTTP-based overlay tunnel, on port 15008.

Figure 3: Two-cluster topology. Each cluster keeps a local InferencePool and picker; the east-west gateways carry HBONE-wrapped mTLS between networks for overflow or failover.
The documented constraints, per Istio’s install guide, matter for an LLM platform (Istio ambient multicluster docs):
- Only multi-network deployments are supported; single-network multicluster is untested and possibly broken, and the Istio beta announcement also describes it as alpha.
- Only multi-primary control planes work; primary-remote does not.
- Ambient east-west gateways carry meshed mTLS traffic only and cannot expose istiod across networks, so a classic east-west gateway is still needed for control-plane communication.
- The design assumes identically named waypoint deployments in every cluster, with configuration synchronised by you.
- Only the local cluster’s service-scope configuration is used as the source of truth, so divergent scope settings across clusters give surprising visibility.
- Remote-network failover does not distribute traffic evenly across endpoints because of HTTP multiplexing and connection pooling.
That last point is the one an LLM operator should underline. Inference connections are long-lived streams. If the mesh pins a small number of multiplexed HBONE connections to a remote gateway, and that gateway favours particular endpoints, a failover event can concentrate the entire spillover on a few GPUs in the surviving cluster.
The Istio announcement for 1.29 also states that peer metadata is exchanged through HBONE baggage headers to make cross-network L7 telemetry accurate, behind the AMBIENT_ENABLE_BAGGAGE flag, and it describes the feature as not ready for production (Istio 1.29 multicluster beta blog). Our reading: treat it as suitable for staging, dev and non-critical inference, and keep a tested manual runbook for anything that carries revenue.
Who chooses the cluster?
This is the architectural gap. GIE’s picker chooses a pod within one pool, and the pool is a set of local pods selected by label. Ambient multicluster makes remote services reachable and can fail over across networks, but it does that at the service level with locality and health, not with KV-cache knowledge. So there are three ways to compose them:
- Independent pools per cluster, mesh failover between them. Each cluster runs its own pool and picker. Traffic normally stays local. When the local pool has no healthy endpoints, the mesh or a global load balancer sends requests to the other cluster’s ingress, where its own picker takes over. This is simple, robust, and the pattern we recommend.
- A global front door choosing the cluster, pickers choosing the pod. A DNS or L7 global balancer picks a region by latency or data residency, then each regional gateway does inference-aware routing. Cluster choice is coarse; pod choice is fine.
- Cross-cluster pools. The GIE project defines an alpha
InferencePoolImportresource, cluster-local and controller-managed, that lets anHTTPRoutetarget a pool exported from another cluster. It has carried an alpha designation since v1.1.0 and may change (InferencePoolImport docs). We found no evidence that Istio implements it; do not plan around it.
For comparison, Google’s GKE multi-cluster Inference Gateway is a generally documented product where a config cluster defines routing and KV-cache usage from all clusters informs decisions, with limits: all clusters in one VPC, at most two clusters per pool because of NEG limits, and no Model Armor support (GKE docs). That shows the ambition; it is a managed, cloud-specific implementation rather than something you get by installing Istio.
Installing the pieces
A minimal Istio install for the inference gateway, taken from Istio’s documentation, is:
kubectl kustomize "github.com/kubernetes-sigs/gateway-api-inference-extension/config/crd?ref=v1.0.1" | kubectl apply -f -
istioctl install --set profile=minimal \
--set values.pilot.env.SUPPORT_GATEWAY_API_INFERENCE_EXTENSION=true \
--set values.pilot.env.ENABLE_GATEWAY_API_INFERENCE_EXTENSION=true -y
The route then references the pool as a backend:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: llm-route
spec:
parentRefs:
- name: inference-gateway
rules:
- backendRefs:
- group: inference.networking.k8s.io
kind: InferencePool
name: vllm-qwen3-32b
Add a DestinationRule for the picker Service to enable TLS. Version numbers above are those in Istio’s documentation at the time of writing; always check the current page, since the Gateway API and inference CRD pins move with each Istio release.
Trade-offs, Gotchas, and What Goes Wrong

Figure 4: Failure handling. Each branch is a decision you must make deliberately: what happens when the picker is gone, when the pool is empty, when it is saturated, and when a whole cluster is unhealthy.
The picker is a new single point of failure
Every request now depends on a gRPC service. The failureMode field on endpointPickerRef decides what happens when it is unreachable. FailOpen lets the gateway fall back to its ordinary endpoint selection, so the service stays up but performance reverts to cache-blind balancing, which under a cache-heavy workload can look like an outage of a different kind (time-to-first-token rising several-fold as caches thrash). FailClose rejects requests, which is safer for strict cost or compliance guarantees but makes the picker a hard dependency. Neither is universally right. Run the picker with at least two replicas, give it a PodDisruptionBudget, and alert on the fallback rate, because a silent FailOpen is easy to miss.
Multiple picker replicas raise their own question: each holds a separate view of in-flight requests, so their decisions are less coordinated than a single instance’s. Some implementations keep a shared or leader-based view; check what yours does before assuming linear scaling.
Stale metrics and herding
Metrics scraped from model servers lag reality. During a traffic burst, every request in the scrape window sees the same “least loaded” pod. The mitigation is to combine scraped state with the picker’s own in-flight counts, and to use weighted-random rather than max-score selection when bursts are common. The llm-d docs list both a max-score-picker (the default) and a weighted-random-picker that uses scores as sampling probabilities. Test with a synthetic burst, not steady load.
Prefix affinity can create hotspots
The very feature that produces headline speedups concentrates load. One viral system prompt means one pod is the favoured destination. Without a load guard, that pod’s queue grows until its queue-depth score overrides the affinity, and you oscillate. Set weights so that queue depth dominates when a pod is past a threshold, and verify with a workload whose prefix distribution is skewed, as real ones are.
Long-lived streams defeat naive failover
LLM responses are streamed over HTTP/2 or server-sent events for many seconds. Two consequences follow. A pod drained for a deploy must wait for streams to finish, so termination grace periods need to be minutes, not seconds. And when a cross-cluster path degrades, existing streams die while new ones move; clients need retry logic that resends the whole request, since a half-generated response cannot be resumed. The ambient multicluster note that remote failover does not spread evenly across endpoints compounds this: after a failure the surviving cluster’s ingress may see a thundering herd of retries that all pass through the same few connections.
Costs of the extra hop
The picker adds a network round trip and CPU per request. For a request that will take seconds to generate, a few milliseconds of scheduling is usually negligible, and cache-aware routing more than pays for itself. For very short completions or embedding calls that take 20 milliseconds, the overhead is proportionally larger and the benefit is smaller. Route those through a plain Service.
Version and status drift
The fastest way to get hurt is to mix documentation from different eras. InferenceModel appeared in early alpha material and Istio’s own blog warned it was likely to change; the project later replaced it with InferenceObjective, then removed the alpha types from its tree at v1.6.0. Documentation conflicts also exist on smaller points: one project page says the endpointPickerRef field became optional in v1.5.0 while the v1.6.0 notes describe making it optional, so read the CRD you actually installed. The GIE releases moved through v1.4, v1.5 (flow control, a pluggable parser framework) and v1.6 within months, and the pace suggests pinning exact versions.
Security and multi-tenancy
An inference gateway sees prompts. If the picker logs bodies for debugging, you have created a store of sensitive text. Restrict its logging, keep the gateway-to-picker channel on TLS (Istio requires a DestinationRule), and remember that prefix-cache-aware routing has a side channel: a tenant who can measure time-to-first-token might infer whether another tenant’s prompt prefix is cached. If tenants must not share cache state, partition pools per tenant or disable prefix-sharing across them. This is an inference from how the mechanism works, not a documented incident.
Anti-patterns
- Treating the mesh as the router: ambient gives identity and encryption, not KV-cache awareness.
- One giant pool spanning heterogeneous GPUs: A100 and H100 replicas have different queue and cache economics. Use separate pools per hardware class.
- Enabling cross-cluster failover without capacity headroom: the surviving cluster must absorb the load or you convert a partial failure into a full one.
- Autoscaling on CPU: use queue depth or KV-cache utilization, not CPU, for GPU replicas.
Practical Recommendations
Start small and prove value in one cluster before adding the mesh dimension. The order of operations that we think minimises risk is as follows.
Phase 1: single-cluster inference gateway. Install the CRDs and the Istio flags, deploy one pool and one picker, and route a shadow slice of traffic. Compare P50 and P95 time-to-first-token and tokens per second per GPU against your existing Service-based path. If your workload has no shared prefixes, expect modest gains and decide whether the added component is worth it.
Phase 2: harden the picker. Two replicas, PodDisruptionBudget, DestinationRule TLS, an explicit failureMode, and dashboards for picker latency, fallback rate, and per-pod queue depth.
Phase 3: second cluster with independent pools. Give each cluster its own pool and picker. Add ambient multicluster only for the failover path and for shared services such as auth or a vector store, in a non-production environment first, given the beta status. Rehearse a full-cluster loss.
Phase 4: revisit cross-cluster pooling when your chosen implementation documents support for exported pools, not before.
Checklist:
- Pin Gateway API, GIE, Istio and vLLM versions together and record them.
- Choose and test
failureModedeliberately. - Alert on picker fallback rate, not just picker availability.
- Load-test with skewed prefixes and with bursts.
- Set termination grace periods to cover your longest generation.
- Keep separate pools per GPU class and per tenant where isolation matters.
- Keep waypoint names and service-scope settings identical across clusters.
- Autoscale on queue depth or KV-cache utilization.
- Write a manual cross-cluster failover runbook and drill it.
Frequently Asked Questions
What is the Gateway API Inference Extension?
It is an official Kubernetes project that turns a Gateway API gateway into an inference gateway. It adds an InferencePool backend type and defines a protocol, built on Envoy’s ext_proc, by which an endpoint picker chooses the best model-server pod per request using metrics such as queue depth, KV-cache use, prefix-cache hits and loaded LoRA adapters. It standardises the API and protocol; the scheduling logic is pluggable.
Does Istio support the Gateway API Inference Extension?
Yes. Istio documents support from release 1.27, enabled by two pilot environment flags and the inference CRDs, and the March 2026 announcement labels it beta. It works at the ingress Gateway through an endpoint picker. Waypoint-level integration was described as future work, and we could not verify it as shipped, so do not assume in-mesh east-west inference routing.
Is Istio ambient multicluster production ready?
Not by the project’s own most cautious wording. It reached beta in Istio 1.29 (February 2026), and the beta announcement says it is not ready for production, while the install guide calls multi-network beta with caveats. Only multi-primary, multi-network topologies are supported. Use it in staging and for non-critical paths, and plan for uneven failover load.
Can one InferencePool span two clusters?
Not in any implementation we could verify for Istio. GIE defines an alpha InferencePoolImport resource for pools exported across clusters, but it may change and we found no Istio support. The dependable pattern is one pool and picker per cluster, with a global balancer or mesh failover between clusters. GKE offers a managed multi-cluster variant with its own limits.
Do I still need a service mesh if I use an inference gateway?
They solve different problems. The inference gateway chooses which GPU pod handles a request. A mesh provides mutual TLS, identity, authorization and telemetry between workloads, and with ambient multicluster, connectivity between networks. You can run the inference gateway without ambient. Add ambient when you need consistent zero-trust policy across the gateway, model servers and other services.
How much does inference-aware routing improve latency?
It depends on prefix reuse. In llm-d’s published benchmark with a cache-friendly workload, P90 time-to-first-token fell from roughly 92.6 seconds under random scheduling to 0.54 seconds. That is a best-case, vendor-run result. Workloads with unique prompts will see far smaller gains, mostly from queue-depth and KV-pressure awareness. Measure on your own traffic with a shadow deployment before committing.
Further Reading
- Istio ambient mesh versus Linkerd: an architecture decision record
- Gateway API 1.6, TCPRoute and UDPRoute versus LoadBalancer
- Kubernetes Gateway API versus Ingress in 2026
- vLLM versus SGLang versus TensorRT-LLM
- Gateway API Inference Extension documentation
- Istio: Gateway API Inference Extension task
By Riju — about

Pingback: llm-d on Kubernetes: Disaggregated LLM Inference