Eleven small language models compared for CPU-only inference in 2026: real Q4 file sizes, official GGUF availability, licences and the memory-bandwidth ceiling.
How to run small language models (SLMs) on-device: model sizing, distillation, quantization, NPU acceleration, memory budgets, and when a 1-8B SLM beats a cloud LLM.