SGLang 0.5.16-0.5.18 add DSpark speculative decoding, a Rust server and Breakable CUDA Graph. Breaking flags, cache moves and a step-by-step upgrade path.
Upgrading vLLM 0.28 to 0.30: Model Runner V2 is now default, MRV1 removal targeted for v0.32, new flags, removed env vars and a safe rollout checklist.
TensorRT 11 removed the entire IPluginV2 family, weak-typing builder flags, and implicit quantization. What breaks in your engine build and how to port it.
Karpenter vs Cluster Autoscaler for GPU node scaling in 2026: bin-packing, spot, cold-start, consolidation and the real cost difference for ML clusters.