TensorRT 11 removed the entire IPluginV2 family, weak-typing builder flags, and implicit quantization. What breaks in your engine build and how to port it.
An AI inference cost optimization decision record: continuous batching, KV-cache, quantization, speculative decoding, spot GPUs, and autoscaling the inference path.