Synthetic data for LLM fine-tuning explained: self-instruct, distillation, rejection sampling, quality filters, decontamination and how to avoid model.
Which fine-tuning method for a small model you will deploy on CPU: VRAM budgets, the merge-target trap worth 14 points, and when distillation is actually worth it.
A 2026 cost and quality decision record for fine-tuning vs RAG vs long-context LLMs: token economics, latency, accuracy trade-offs, and a decision matrix.