SGLang 0.5.16-0.5.18 add DSpark speculative decoding, a Rust server and Breakable CUDA Graph. Breaking flags, cache moves and a step-by-step upgrade path.
The 500ml and 2.9 Wh figures come from real papers that said something narrower. Here is every published measurement of AI energy and water per query, with its scope.
MLPerf Inference v6.1 added an Edge Agentic benchmark. NVFP4, FP8 KV cache, 96% cache reuse and tree MTP explain the 6.4x gap over the llama.cpp reference.
Diffusion LLMs explained: how text diffusion models replace next-token prediction with iterative denoising and parallel decoding - masked discrete diffusion, remasking, throughput trade-offs vs autoregressive models in 2026.
How constrained decoding guarantees valid LLM output: grammars, FSAs, token masking, JSON-schema enforcement, and where structured generation breaks in production.
A 2026 benchmark methodology for small language models on edge GPUs — latency, tokens/sec, memory, and cost for Phi, Gemma, and Qwen on Jetson-class hardware.