How Istio ambient multicluster and the Gateway API Inference Extension route LLM traffic by KV-cache and queue depth, with a reference architecture and failure modes.
SGLang 0.5.16-0.5.18 add DSpark speculative decoding, a Rust server and Breakable CUDA Graph. Breaking flags, cache moves and a step-by-step upgrade path.
Upgrading vLLM 0.28 to 0.30: Model Runner V2 is now default, MRV1 removal targeted for v0.32, new flags, removed env vars and a safe rollout checklist.
An AI inference cost optimization decision record: continuous batching, KV-cache, quantization, speculative decoding, spot GPUs, and autoscaling the inference path.