SGLang 0.5.16-0.5.18 add DSpark speculative decoding, a Rust server and Breakable CUDA Graph. Breaking flags, cache moves and a step-by-step upgrade path.
Q2 2026 LLM inference benchmark across vLLM, TGI, SGLang, and Triton — throughput, p50/p99 TTFT/TPOT, KV-cache efficiency, and which engine wins per workload class.