llm-d on Kubernetes: Prefill/Decode Disaggregation for LLM Inference Posted by By MPRAUTO MPRAUTO September 29, 2026Posted inKubernetesNo Comments llm-d explained: disaggregated prefill and decode, KV-cache-aware routing and vLLM on Kubernetes, with cost math, an install path and where it beats a plain vLLM deployment.