Synthetic Data for LLM Fine-Tuning: Generation, Quality Filtering and Avoiding Model Collapse
Most teams that fine-tune a language model discover the same bottleneck within a week: the model is fine, the training code is fine, and the labelled data does not exist. Human-written instruction and response pairs are slow, expensive and inconsistent. So the data gets generated by another model, and the real work begins, because a naive generation loop produces fluent, repetitive, subtly wrong text that makes the student worse. Synthetic data for LLM fine-tuning is now standard practice, yet the difference between a pipeline that lifts a model and one that quietly degrades it is almost entirely in the filtering, diversity control and evaluation hygiene around the generator.
This matters now because open-weight models are strong enough to act as teachers, tooling for verification is mature, and the research on what goes wrong is no longer speculative. You can reason about it from published results instead of folklore.
You will leave with a working mental model of the whole pipeline: how to generate with Self-Instruct, Evol-Instruct and persona prompts, how to filter with verifiers, judges and MinHash, how to decontaminate against evaluation sets, how to stay clear of model collapse, and a runnable Python skeleton you can adapt.
What this covers: generation methods, distillation and rejection sampling, quality and diversity filters, decontamination, the model-collapse literature, licensing limits on teacher outputs, illustrative cost arithmetic, a code skeleton and a decision matrix.
Context and Background
Instruction tuning turned base language models into assistants, and for the first generation of open models the instruction data came from people. The turning point for synthetic data was Self-Instruct (Wang et al., 2022), which showed that a model could bootstrap its own instruction set from a small human-written seed. In the authors’ released repository the seed is 175 tasks and the output is a dataset of 52 thousand instructions paired with 82 thousand instance inputs and outputs. Applied to vanilla GPT-3, the paper reports a 33 percent absolute improvement on Super-NaturalInstructions, comparable to InstructGPT-001, while remaining about 5 points behind it in a human evaluation on novel expert-written tasks. The point was not that the data was perfect. It was that noisy generated data, filtered sensibly, was good enough to move a model most of the way.
What followed was a rapid cycle of refinements. WizardLM’s Evol-Instruct (Xu et al., 2023) used an LLM to rewrite seed instructions into progressively more complex ones, and reported that its evolved data beat human-written instructions in human evaluation on the authors’ test sets. Persona Hub (Ge et al., 2024) attacked diversity head-on with a collection of one billion personas curated from web data, used to steer a generator towards different perspectives. Distillation from strong proprietary and open teachers became the dominant recipe for small models. We cover the training-side choices in our guide to LoRA, QLoRA, full fine-tuning and distillation for small language models; this post is about the data those methods consume.
There is also a counter-current. In 2024 Shumailov and colleagues published “AI models collapse when trained on recursively generated data” in Nature (volume 631, pages 755 to 759). It showed that training successive generations of models on the output of their predecessors degrades them, with the rare parts of the distribution disappearing first. That paper is regularly cited as proof that synthetic data is poisonous. It is not, and reading it precisely matters, because the setup it studies, replacing real data entirely with model output generation after generation, is not what a competent fine-tuning pipeline does.
Finally, synthetic data is a decision with alternatives. If your gap is knowledge rather than behaviour, retrieval is usually cheaper than any training run, which our comparison of fine-tuning versus RAG versus long context works through. Synthetic data earns its place when you need to change how the model behaves, formats, reasons or refuses, not merely what documents it can see.
A note on sourcing. The facts attributed to papers in this article come from the papers’ own abstracts and repositories, which we checked while writing. Cost and throughput figures are illustrative arithmetic, labelled as such, because real prices change monthly and vary by provider.
The Reference Pipeline: Generate, Filter, Mix, Evaluate
Direct answer: A sound synthetic-data pipeline has five stages: seed with real human data, generate candidates with a teacher, filter them for duplicates and quality, remove anything overlapping your evaluation sets, and mix the survivors with real data before fine-tuning. Evaluation on held-out real data closes the loop and decides whether the next round is worth running.

Figure 1: End-to-end pipeline for synthetic data for LLM fine-tuning, with the evaluation loop feeding back into generation.
The diagram shows the stages as a line, but the economics are a funnel. A typical run generates many more candidates than it keeps, and the ratio is a design parameter. Deduplication usually removes the most obvious waste first, verification removes the wrong, and decontamination removes the dangerous. What remains is a small, clean set that should outperform a much larger unfiltered one. The evaluation arrow back to the generator is what makes this engineering rather than hope: if held-out real performance does not rise, you change the seeds, prompts or filters, not the learning rate.
Seed data is the leverage point
Every method below amplifies whatever you feed it. A generator seeded with 175 careful, varied tasks produces varied output; seeded with 20 near-identical examples it produces 52 thousand near-identical examples with different nouns. Seed quality has an outsized effect because the model learns the distribution of the seeds and then extrapolates it, so gaps in the seeds become gaps in the data.
Treat seeds as the one place where human time is worth spending. Write them to cover the task types, formats, lengths, difficulty levels and refusal behaviours you want. Include hard cases and edge cases, not only the happy path. If the production workload has a long tail, such as unusual units, rare languages or malformed inputs, put representatives of that tail in the seeds explicitly, because the generator will not invent them on its own.
The generator decides the ceiling
The teacher sets an upper bound on quality, with a nuance. A student can exceed a noisy teacher when a verifier selects only the correct outputs, which is the idea behind rejection sampling and why it is so valuable for math and code. Without a verifier, the student mostly inherits the teacher’s strengths and its errors, plus a compression loss. Pick the strongest teacher you are legally and economically able to use, and spend the saved effort on verification.
Teacher choice also shapes style. Models trained on one family’s output absorb its phrasing, its hedging habits and its characteristic mistakes. Mixing teachers is a cheap diversity lever and reduces the risk of overfitting to a single model’s quirks.
Quality, diversity and volume trade against each other
Three properties matter and you cannot maximise all of them for free. Quality means the response is correct and well formed. Diversity means the prompts and responses cover many different situations. Volume means enough examples for the model to learn the pattern. Cheap generation inflates volume at the expense of diversity, aggressive filtering raises quality at the expense of volume, and over-evolving prompts raises difficulty at the expense of realism. A useful rule from practice: when you must choose, a few thousand varied, verified examples beat a few hundred thousand repetitive ones for narrow behavioural fine-tuning.
Generation Methods in Detail
Each method below answers a different question. Self-Instruct asks how to get breadth from a small seed. Evol-Instruct asks how to get difficulty. Persona prompting asks how to get diversity. Distillation asks how to transfer a capability. Rejection sampling asks how to get correctness. Real pipelines combine several.
Self-Instruct: bootstrap breadth from a small seed
Self-Instruct runs a loop. Sample a few instructions from the pool, show them to the model as in-context examples, and ask for new instructions. Filter the new ones, generate inputs and outputs for those that survive, add them to the pool, and repeat. The original filters were deliberately simple: discard instructions too similar to existing ones using ROUGE-L overlap, discard those containing words that signal tasks a text model cannot do, and discard malformed outputs. The loop is shown in Figure 2.

Figure 2: The self-instruct loop. Generated instructions that pass similarity and format filters rejoin the pool and seed the next round.
The mechanism has two failure modes that every reproduction hits. First, the pool drifts towards the model’s comfort zone: instructions get shorter, more generic and more similar, because those are the likeliest completions. The similarity filter slows this but does not stop it. Second, outputs are produced by the same model that wrote the instruction, so there is no independent check on correctness. Self-Instruct works because the base task was loosely specified and noisy data still teaches the instruction-following format. For tasks with a right answer, it is not enough alone.
Evol-Instruct: manufacture difficulty
Evol-Instruct rewrites an existing instruction into a harder or more specific one. In-depth evolution adds constraints, requires more reasoning steps, deepens the topic or concretises an abstract request. In-breadth evolution creates a new instruction in the same domain but a different niche. After each evolution a filter discards failures, such as responses that merely copy the prompt or evolved instructions the model refuses. The WizardLM paper reports that fine-tuning LLaMA on the evolved data gave a model preferred over ChatGPT on the high-complexity portion of its test set, and over 90 percent of ChatGPT’s capability on 17 of 29 skills in GPT-4 evaluation. These are the authors’ numbers on their own test sets with an LLM judge, and should be read as encouraging rather than definitive.
The practical lesson is that difficulty is a dial. If your seed set is all easy, a student trained on it will be confident and brittle. Two or three evolution rounds give a difficulty spread; more than that tends to drift into contrived instructions no user would write. Measure realism by sampling evolved prompts and asking whether a real user of your product could plausibly have typed them.
Persona-driven synthesis: diversity by construction
Rather than hoping the generator wanders into new territory, you can force it there. Persona Hub assembles a large set of personas, one billion in the paper, and conditions each generation on one: write a maths problem a high-school teacher would pose, a coding question a hospital data engineer would ask, a support ticket from a logistics dispatcher. The persona acts as a high-entropy random variable injected into the prompt, and the generator maps it to different content. The paper demonstrates breadth across math and reasoning problems, user instructions, knowledge-rich text, game characters and tool definitions, and it is labelled work in progress; it does not offer a controlled benchmark that proves downstream gains in every setting.
You do not need a billion personas. For a domain product, a few hundred hand-written roles crossed with task types, difficulty levels and formats gives a combinatorial grid that is easy to audit. Sampling uniformly from the grid and tracking coverage counts is a more reliable diversity guarantee than prompt wording.
Distillation: teach a student from a stronger teacher
Distillation in the LLM setting usually means sequence-level distillation: the teacher generates responses, often with reasoning traces, and the student is fine-tuned on them with ordinary supervised learning. It does not require access to teacher logits, which is why it works through an API. When logits are available, token-level distillation with a KL loss transfers more information per example, but most practical pipelines use text only.
Reasoning traces deserve a note. Training a student on a teacher’s chain of thought transfers the reasoning pattern, but the trace can be wrong while the final answer is right. Verifying the final answer, and where possible the steps, avoids teaching the student plausible-looking invalid logic.
Rejection sampling with verifiers
Rejection sampling is the highest-leverage technique when correctness can be checked automatically. Sample k responses per prompt, run each through a verifier, and keep only those that pass. For code the verifier is a test suite or a type checker. For maths it is an exact answer match or a symbolic check. For structured output it is a schema validator. For retrieval-grounded answers it is a check that each claim is supported by the cited passage.

Figure 3: Rejection sampling. Several samples per prompt are scored by a verifier and a judge or reward model, with passing responses kept for supervised fine-tuning and failures optionally reused as negatives.
The economics are favourable because verification is usually far cheaper than generation, and because failures are not waste: a failed sample paired with a passing sample for the same prompt is a ready-made preference pair for DPO-style training. Pass rate also tells you something. If a prompt passes in 0 of 16 samples it is probably beyond the teacher or malformed. If it passes 16 of 16 it teaches little. The informative prompts sit in between, and prioritising them is a cheap form of curriculum.
Quality Filtering: Dedup, Judges and Reward Models
Generation is the easy half. Filtering is what separates a dataset from a pile.
Near-duplicate removal with MinHash
Generators repeat themselves. Left unchecked, thousands of prompts begin with the same phrasing and describe the same scenario. Exact-match deduplication catches almost none of this. MinHash with locality-sensitive hashing (LSH) approximates Jaccard similarity between sets of shingles in sub-quadratic time. Each document is reduced to a short signature; documents whose signatures collide in at least one band become candidate duplicates; candidates can be verified exactly if needed.
Two parameters control behaviour. The number of permutations sets the accuracy of the similarity estimate, and 128 is a common starting point. The Jaccard threshold sets how similar two items must be to count as duplicates. For instruction prompts a threshold around 0.7 to 0.8 on word five-gram shingles is a reasonable starting range, but it should be tuned by inspecting borderline pairs, because short prompts have few shingles and similarity swings widely. Deduplicate on the prompt rather than only on the response, since the diversity you want is in what users ask.
LLM-as-judge scoring
A judge model scores each example on a rubric: correctness, helpfulness, format compliance, safety, and whether the instruction is clear. Judges scale well and catch things rules cannot, but they have known biases. They tend to prefer longer answers, favour outputs from their own model family, and are sensitive to the order in which candidates are shown. Mitigate by using a different model from the generator, randomising order in pairwise comparisons, scoring against explicit rubrics with anchored examples, and calibrating against a few hundred human-labelled items so you know the judge’s true agreement rate before trusting its threshold.
A cheap and effective pattern is a two-stage filter. A fast rule layer removes the obviously broken, such as empty responses, truncated outputs, language mismatches and refusal boilerplate. Only survivors reach the expensive judge. This typically cuts judge cost substantially because most waste is easy to spot.
Reward-model filtering
A reward model trained on human preferences gives a scalar score without a prompt-engineered rubric. It is cheaper per item than a large judge and deterministic, which helps reproducibility. Its weakness is distribution shift: a reward model trained on general chat can mis-score specialised domains, and scores from different models are not comparable. Use it to rank within a prompt, choosing the best of k, rather than as an absolute quality threshold across the dataset.
Diversity and coverage metrics
Quality filters keep good items; they do not guarantee the set is varied. Track diversity separately. Useful cheap signals include the distribution of prompt lengths, the number of distinct first-three-word openings, embedding-cluster counts with per-cluster sizes, and the share of data in the largest cluster. If one cluster holds a large share of the dataset, downsample it. This step often delivers more benefit than any additional quality scoring.
A runnable pipeline skeleton
The skeleton below implements the structural pieces: MinHash dedup, n-gram decontamination, rejection sampling with a pluggable verifier, and a cap on the synthetic share when mixing with real data. The teacher call is stubbed so it runs offline; replace sample_fn with your API or local inference call. It requires pip install datasketch.
import hashlib, random, re
from dataclasses import dataclass
from datasketch import MinHash, MinHashLSH
@dataclass
class Example:
prompt: str
response: str
score: float = 0.0
source: str = "synthetic"
def shingles(text, n=5):
toks = re.findall(r"\w+", text.lower())
return {" ".join(toks[i:i + n]) for i in range(max(1, len(toks) - n + 1))}
def minhash(text, num_perm=128):
m = MinHash(num_perm=num_perm)
for s in shingles(text):
m.update(s.encode("utf8"))
return m
def dedup(examples, threshold=0.7, num_perm=128):
lsh = MinHashLSH(threshold=threshold, num_perm=num_perm)
kept = []
for i, ex in enumerate(examples):
mh = minhash(ex.prompt, num_perm)
if lsh.query(mh): # a near-duplicate is already kept
continue
lsh.insert(str(i), mh)
kept.append(ex)
return kept
def decontaminate(examples, eval_texts, n=13):
banned = set()
for t in eval_texts:
toks = re.findall(r"\w+", t.lower())
for i in range(len(toks) - n + 1):
banned.add(hashlib.md5(" ".join(toks[i:i + n]).encode()).hexdigest())
clean = []
for ex in examples:
toks = re.findall(r"\w+", (ex.prompt + " " + ex.response).lower())
hit = any(hashlib.md5(" ".join(toks[i:i + n]).encode()).hexdigest() in banned
for i in range(len(toks) - n + 1))
if not hit:
clean.append(ex)
return clean
def rejection_sample(prompt, sample_fn, verify_fn, k=8, keep=1):
cands = [sample_fn(prompt) for _ in range(k)]
scored = [(verify_fn(prompt, c), c) for c in cands]
passing = sorted([s for s in scored if s[0] > 0], key=lambda s: -s[0])
return [Example(prompt, c, sc) for sc, c in passing[:keep]]
def mix(real, synthetic, max_synth_ratio=0.5, seed=0):
rng = random.Random(seed)
cap = int(len(real) * max_synth_ratio / (1 - max_synth_ratio))
synth = rng.sample(synthetic, min(cap, len(synthetic)))
out = real + synth
rng.shuffle(out)
return out
We ran this skeleton against a toy arithmetic teacher that answers incorrectly some of the time. The verifier kept only correct answers, dedup removed repeated prompts, and the decontamination pass removed an item that overlapped a planted evaluation string. That confirms the plumbing, not the quality of any real dataset. Note the design choices: the LSH index is queried before insertion so the first occurrence survives; decontamination checks both prompt and response; and mix bounds the synthetic share relative to real data, which becomes important in the model-collapse section.
Decontamination: Keeping Evaluation Sets Out of Training
Data decontamination is the step teams skip and regret. If your synthetic data contains the questions, or close paraphrases of the questions, in your evaluation set, your benchmark scores measure memory, not ability. Synthetic data is unusually exposed to this because the teacher may have seen public benchmarks during its own training and will happily reproduce them when asked for “a challenging maths problem”. Contamination therefore enters through the teacher even if you never touched a benchmark.
What to compare against
Build a single registry of every evaluation set you care about: public benchmarks, internal held-out sets, customer acceptance tests and red-team prompts. Compare synthetic prompts and responses against all of them. Include the answers, not only the questions, since a leaked answer string is also a signal.
Methods, from cheap to thorough
The standard cheap method is long n-gram overlap: if a training item shares a run of, say, thirteen consecutive tokens with an evaluation item, drop the training item. The skeleton above does this with hashed n-grams. Published model reports have used n-gram schemes of this kind with different n values, and the right n is a trade-off: small n produces false positives on common phrases, large n misses light paraphrases.
Paraphrase detection needs more. Embed both sets and flag training items whose nearest evaluation neighbour exceeds a cosine similarity threshold, then review the flagged pairs by hand or with a judge. This catches reworded versions at the cost of false positives. For code, normalise identifiers and whitespace before comparing, or compare abstract syntax trees, because a renamed variable defeats string matching.
Finally, decontaminate more than once. Run it after generation, again after any evolution round, and again at the final mix, since each stage can reintroduce overlap. Log how many items each pass removes. A sudden jump in removals after changing the teacher or seed is itself a useful alarm.
Contamination is not only about benchmarks
The same hygiene applies to personal and proprietary data. If seeds or retrieval contexts contain customer records, a generator can reproduce them verbatim in the synthetic output. Run PII detection on the synthetic set as well as on the seeds, and keep seeds out of the training data unless you intend them there.
Model Collapse: What the Literature Actually Shows
Model collapse is the most misread result in this area, so precision matters.
Shumailov et al. (Nature, 2024) studied what happens when each generation of a model is trained on data produced by the previous generation. They report that the process causes irreversible defects: the tails of the original content distribution disappear first, a stage the authors call early collapse, and later generations converge to a distribution with little resemblance to the original and much lower variance, which they call late collapse. They show the effect in language models, variational autoencoders and Gaussian mixture models, with fine-tuned OPT-125m models showing rising perplexity and increasingly implausible text over successive generations. Penalising repetition did not eliminate the effect in their experiments.
The mechanism is statistical. A finite sample from a distribution under-represents rare events. A model fitted to that sample under-represents them further. Sample from that model and fit again, and the rare events thin out until they vanish. Because the learned function is also approximate, errors compound across generations. It is a feedback loop with a systematic bias towards the mode.

Figure 4: Replacing real data with each generation’s output loses the tails and collapses variance; keeping real data and accumulating synthetic data stays stable.
The key qualifier: replace versus accumulate
Gerstgrasser et al. (2024), in “Is Model Collapse Inevitable?”, examined the assumption hidden in the collapse experiments: that each generation’s synthetic data replaces the previous data. When the original real data is retained and each generation’s synthetic output is added to it, they report that collapse is avoided across the model sizes, architectures and hyperparameters they tested, with similar findings for diffusion models and variational autoencoders. In a linear-model framework they prove that test error under accumulation stays bounded as iterations increase, whereas under replacement it grows.
Shumailov and colleagues themselves conclude that access to genuine human-generated data remains valuable. These two results are compatible. The practical reading: model collapse is a risk of recursive replacement, not of using synthetic data as such. A fine-tuning pipeline that begins from a pretrained model, mixes in real examples, and uses a stronger or independent teacher is not running the experiment that collapses models. A pipeline that trains model N+1 only on the filtered output of model N, repeatedly, and discards real data is.
Practical guardrails against collapse
First, always keep a real-data anchor in every training mix, and never drop it between rounds. Second, bound the synthetic fraction and ablate it: run the same fine-tune at 0, 25, 50 and 75 percent synthetic share and plot held-out real performance. Third, prefer an external or stronger teacher over the student’s own previous output. Fourth, watch the tails explicitly: track performance on rare categories, long inputs and minority languages, since those degrade before aggregate metrics move. Fifth, track diversity metrics round over round; falling entropy in prompts or responses is the early warning. Our rule of thumb is that if you cannot name the real data in your mix, you are not mixing.
There is a related effect worth separating from collapse: sycophantic or stylistic homogenisation. A student trained on a single teacher’s output inherits its register, and a team deploying several students from the same teacher ends up with similar blind spots. This is not the mathematical collapse in the paper, but it is a real product risk, and the mitigations overlap: multiple teachers, real data and measured diversity.
Licensing and Terms of Service for Teacher Outputs
Legal constraints decide which teacher you may use, and they are not uniform. This is an engineering summary, not legal advice; read the current terms yourself and ask counsel for anything commercial.
Some providers restrict using their outputs to build competing models. OpenAI’s terms of use, for example, list among prohibited activities “Use Output to develop models that compete with OpenAI.” Other providers have comparable clauses with different wording and scope. Open-weight licences differ again: some permit using outputs to improve other models, sometimes with attribution or naming conditions, while others restrict commercial use or impose thresholds. Llama community licences, for instance, have included conditions of this kind that changed between versions, so check the exact text of the version you use rather than relying on memory.
Three practices keep you out of trouble. Record, per example, which teacher and which version produced it, so you can remove a teacher’s data if its terms change or a dispute arises. Prefer teachers whose licences clearly permit derivative training, such as permissively licensed open-weight models, for any dataset you plan to release or commercialise. And treat the dataset you publish as a separate legal object from the model you train: releasing outputs from a restricted teacher can breach terms even if the trained model is private.
Provenance also helps technically. A source field on every example, as in the skeleton, lets you slice evaluation by teacher and find which source helps or hurts.
Cost Arithmetic (Illustrative)
All numbers below are illustrative assumptions for planning, not quotes. Substitute your own prices and measured rates.
Suppose you want 20,000 verified training examples. Assume an average of 400 prompt tokens and 600 response tokens per sample, eight samples per prompt for rejection sampling, and a verifier pass rate of 40 percent after filtering. You need about 50,000 prompts to yield 20,000 verified examples at roughly one kept response per passing prompt, so the generation volume is 50,000 prompts times 8 samples times 1,000 tokens, or 400 million tokens. At an assumed blended price of 1 US dollar per million tokens for an open-weight teacher served on your own GPUs or a cheap API, that is about 400 dollars of generation. At an assumed 10 dollars per million for a frontier API it is about 4,000 dollars.
Judge passes add cost. If a judge reads each surviving sample at 1,200 tokens, and half the 400,000 samples pass the rule layer, that is roughly 240 million judge input tokens, so the judge may cost as much as generation. This is why the rule layer and a cheaper judge matter. Compare against the human alternative: if an annotator produces, say, 30 verified items an hour and costs 15 dollars an hour, 20,000 items cost 10,000 dollars before review overhead. Synthetic data is cheaper under these assumptions, but the saving shrinks fast if you need heavy human review of the output, which for safety-critical domains you will.
The sensitivity worth noticing is the pass rate. Halving it doubles the generation bill, and a weak verifier that passes too much shifts cost into bad data instead. Measure pass rate on a pilot of a few hundred prompts before committing to a full run.
Choosing a Method: Decision Matrix
| Goal | Best-fit method | Needs verifier? | Main risk | Typical filter |
|---|---|---|---|---|
| Teach instruction-following format | Self-Instruct style bootstrap | No | Drift to generic prompts | ROUGE or MinHash dedup |
| Raise difficulty and reasoning depth | Evol-Instruct plus rejection sampling | Helpful | Contrived prompts | Judge plus answer check |
| Cover many user types and domains | Persona or grid-driven generation | No | Superficial variety | Cluster balancing |
| Code, math, structured output | Rejection sampling with verifier | Yes | Narrow test coverage | Tests and schema checks |
| Transfer a strong model’s skills | Distillation from teacher traces | Preferred | Licence limits, copied errors | Final-answer verification |
| Improve alignment and tone | Preference pairs from pass and fail samples | Judge or reward model | Judge bias | Position-randomised judging |
If you can verify automatically, start with rejection sampling; it is the only row where the student can plausibly exceed the teacher on the verified slice. If you cannot verify, invest in seeds, diversity control and a calibrated judge, and keep expectations modest.
Trade-offs, Gotchas, and What Goes Wrong
The most common failure is a quiet one: held-out scores on an in-distribution synthetic test set rise while real-world quality does not. The synthetic test set shares the generator’s biases, so the student learns to please the generator. Always evaluate on real, human-written data that never touched the pipeline.
Verifiers can be gamed. A code verifier that checks only that tests pass rewards solutions that special-case the tests. A maths verifier that matches a final number accepts wrong reasoning that lands on the right answer. Audit a sample of passing items manually every round, and strengthen verifiers where you find leaks.
Judges drift with prompt wording, and thresholds tuned for one batch misbehave on the next. Freeze the judge model version and rubric, store its raw scores, and recalibrate against human labels when either changes.
Over-filtering is its own failure. A pipeline that keeps only the highest-scoring items converges on one safe style and removes hard, unusual cases the model most needs. Keep a stratified sample of borderline items and look at what you are rejecting.
Refusal and safety behaviour can regress. Fine-tuning on helpful-sounding synthetic data without safety examples can erode refusals, and inheriting a permissive teacher’s behaviour can import its gaps. Include safety cases in the seeds and evaluate them explicitly.
Finally, reproducibility. Record the teacher version, sampling temperature, prompts, seeds, filter thresholds and library versions. Without that, you cannot explain why round three beat round four.
Practical Recommendations
Begin small and measure. Hand-write a few hundred varied seeds, generate a pilot of a few thousand examples, filter them, and fine-tune a small model to see whether held-out real performance moves. Only scale generation after the pilot shows a gain, because scale amplifies whatever flaws the pilot has.
Spend your engineering effort on verification and on evaluation. Verifiers turn a noisy generator into a reliable source for code, maths and structured tasks, and a real-data held-out set tells you whether any of it worked. Keep a real-data anchor in every mix, bound and ablate the synthetic share, and prefer multiple or external teachers over recursive self-training.
Record provenance from day one. Source, teacher version, licence and filter decisions on every example make audits, removals and ablations cheap, and make the legal conversation easy.
- Write 100 to 500 varied seeds including hard cases and refusals.
- Generate with at least two teachers or two prompt styles.
- Deduplicate prompts with MinHash and review borderline pairs.
- Verify automatically wherever a check exists; calibrate the judge elsewhere.
- Decontaminate after every generation or evolution step.
- Keep real data in every training mix; ablate synthetic share at 0, 25, 50 and 75 percent.
- Evaluate on held-out real data and on rare-category slices.
- Log teacher, version, licence and thresholds per example.
For the choice between training and retrieval, revisit the fine-tuning versus RAG versus long-context comparison, and for how to run the resulting fine-tune cheaply see the LoRA, QLoRA and distillation guide.
Frequently Asked Questions
Is synthetic data good enough to fine-tune an LLM?
Yes, for many tasks, provided it is filtered and mixed with real data. Self-Instruct showed that filtered model-generated instructions could lift a base model most of the way to an instruction-tuned one, and rejection sampling with verifiers can produce high-quality data for code and maths. Quality depends on seeds, verification and diversity. Evaluate on human-written held-out data to confirm that the gain is real and not an artefact of the generator.
What is model collapse and does it affect fine-tuning?
Model collapse is the degradation that occurs when models are trained recursively on their own output, losing rare patterns first and then variance, as shown by Shumailov et al. in Nature in 2024. It mainly threatens pipelines that replace real data with synthetic data over many generations. Work by Gerstgrasser et al. reports that accumulating synthetic data alongside retained real data avoids it. Keep real data in every mix and you avoid the failure mode.
How much synthetic data should I mix with real data?
There is no universal ratio; it depends on task, teacher quality and verification. Run an ablation at 0, 25, 50 and 75 percent synthetic share and compare held-out real performance, including rare slices. Start conservatively, such as one synthetic example per real one, and increase only while real-data metrics keep improving. If gains plateau or tail categories degrade, you have passed the useful point.
How do I filter low-quality synthetic data?
Use layers. First rules remove empty, truncated, wrong-language and refusal outputs. Then MinHash removes near-duplicate prompts. Then automatic verifiers check correctness where possible, and an LLM judge or reward model scores the rest against a rubric, ideally a different model from the generator. Finally, check diversity with clustering and downsample over-represented clusters. Calibrate any judge against a few hundred human labels before trusting its threshold.
Can I train on outputs from GPT-class or other proprietary models?
It depends on the provider’s current terms. Some prohibit using outputs to develop competing models; OpenAI’s terms of use list that as a prohibited activity. Open-weight licences vary and may allow it with conditions. Check the exact terms for the version you use, record the teacher on every example, and take legal advice before commercialising or releasing a dataset derived from a restricted teacher.
What is data decontamination and why does it matter?
Data decontamination removes training examples that overlap with evaluation sets, so benchmark scores measure ability instead of memorisation. Synthetic pipelines are exposed because the teacher may have seen public benchmarks and reproduce them. Use long n-gram overlap for exact matches, embedding similarity for paraphrases, and rerun after each generation stage. Log removal counts; sudden changes after switching teacher or seed usually indicate a contamination source worth investigating.
Further Reading
Internal:
- LoRA vs QLoRA vs full fine-tuning vs distillation for small language models: choosing the training method that will consume your synthetic dataset.
- Fine-tuning vs RAG vs long context in 2026: deciding whether you need training data at all.
External primary sources:
- I. Shumailov et al., “AI models collapse when trained on recursively generated data”, Nature 631, 755 to 759, 2024.
- Y. Wang et al., “Self-Instruct: Aligning Language Models with Self-Generated Instructions”, 2022, and the released code and data.
- M. Gerstgrasser et al., “Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data”, 2024.
- C. Xu et al., “WizardLM: Empowering large pre-trained language models to follow complex instructions”, and T. Ge et al., “Scaling Synthetic Data Creation with 1,000,000,000 Personas”.
This article is educational engineering analysis, not legal advice. Cost figures are illustrative assumptions, not provider quotes.
By Riju — about
