ONNX Runtime 1.30 adds Arm NEON/SVE LinearAttention, INT4 paged KV cache, Go bindings and FP16 fallback changes. What changed and how to upgrade safely.
MLPerf Inference v6.1 added an Edge Agentic benchmark. NVFP4, FP8 KV cache, 96% cache reuse and tree MTP explain the 6.4x gap over the llama.cpp reference.
Eleven small language models compared for CPU-only inference in 2026: real Q4 file sizes, official GGUF availability, licences and the memory-bandwidth ceiling.
How to run small language models (SLMs) on-device: model sizing, distillation, quantization, NPU acceleration, memory budgets, and when a 1-8B SLM beats a cloud LLM.
An AI inference cost optimization decision record: continuous batching, KV-cache, quantization, speculative decoding, spot GPUs, and autoscaling the inference path.
Step-by-step guide to running ML models on ESP32 using TensorFlow Lite Micro — quantization, memory budgeting, ESP-NN acceleration, and deployment patterns.
Edge AI inference at scale, updated for 2026: NVIDIA Jetson Thor, Hailo and Arm Ethos NPUs, INT4/FP8 quantization, runtimes, and how to pick edge accelerators by TOPS-per-watt.