Which fine-tuning method for a small model you will deploy on CPU: VRAM budgets, the merge-target trap worth 14 points, and when distillation is actually worth it.
Eleven small language models compared for CPU-only inference in 2026: real Q4 file sizes, official GGUF availability, licences and the memory-bandwidth ceiling.