DPO vs RLHF vs SFT: LLM Alignment Methods Compared (2026)
Every team that post-trains a language model eventually faces the same question: do we imitate good answers, learn from preferences, or run reinforcement learning against a reward? The DPO vs RLHF debate dominated 2023 and 2024, and it was usually framed as “simple versus powerful.” By 2026 that framing is out of date, because the frontier moved again. Reasoning models are trained with reinforcement learning against programmatic verifiers, and the open recipes that matter combine three or four methods in sequence rather than picking one.
This post rebuilds the comparison from the losses up. We derive what each objective actually optimizes, say what data and compute each one needs, explain why some of them fail in characteristic ways, and map them to what published recipes really use. Nothing here depends on a single benchmark table, because the honest finding in the literature is that rankings flip with the setup.
What this covers: the alignment problem and where each method sits in the pipeline, the math of SFT, the reward-model-plus-PPO loop, and DPO, then IPO, KTO, ORPO, SimPO, GRPO and RLVR, a cost and stability comparison, failure modes, and a decision procedure you can apply to your own project.
Context and Background
Pretraining produces a model that predicts the next token of internet-scale text. That objective rewards fluency and coverage, not helpfulness or safety, so a raw base model will happily continue a harmful request or ramble past the answer. Post-training is the family of techniques that reshapes the model after pretraining so that it follows instructions, declines unsafe requests, and stays calibrated about what it knows.
The canonical reference is the InstructGPT paper from OpenAI, which combined supervised learning on labeler demonstrations with reinforcement learning from human feedback on ranked outputs. Its headline result, reported in the InstructGPT paper, was that outputs from a 1.3B parameter InstructGPT model were preferred to those of the 175B parameter GPT-3, despite having roughly 100 times fewer parameters. That result established the three-stage recipe (SFT, reward model, PPO) as the default for years.
That recipe is expensive and fragile. It needs a reward model, a reference model, a value network, and a live policy in memory at once, and it requires fresh samples from the policy at every step. In 2023 Rafailov and colleagues showed that for the standard KL-regularized preference objective you can skip the reward model and the RL loop entirely. Their Direct Preference Optimization paper reported that DPO exceeded PPO-based RLHF in controlling the sentiment of generations and matched or improved response quality in summarization and dialogue, while being substantially simpler to implement and train.
Then the picture got more complicated. Later work argued that DPO has fundamental limitations and that a well-tuned PPO beats it, that simpler REINFORCE-style estimators match PPO at lower cost, and that for verifiable domains such as math and code, reinforcement learning against a checker beats anything built on human preferences. Open recipes such as Tulu 3 now chain SFT, DPO and a verifiable-reward RL stage. If you are weighing these options for a small model, our guide to LoRA vs QLoRA vs full fine-tuning vs distillation covers the parameter-efficiency side that sits underneath all of them.
What “alignment” means operationally
In practice, alignment is not a single property. Teams optimize for instruction following, refusal behavior on unsafe prompts, tone and format, factuality and calibration, and, increasingly, correctness on tasks with checkable answers. These targets have different signal types. Instruction following can be taught by demonstration. Tone and refusal boundaries are comparative judgments, which suit preference data. Correct math is binary and verifiable, which suits programmatic rewards.
That mapping from target to signal type is the most useful organizing idea in this whole topic. Each method in the sections below is best understood as a different way of converting one kind of signal into gradient updates on the policy.
Where Each Method Sits: The Alignment Pipeline
Modern post-training is a pipeline, and the three headline methods are stages rather than rivals. SFT almost always comes first because every other method needs a policy that already produces reasonable answers. After SFT, the signal you have decides the next step.

Figure 1: The post-training decision point. SFT produces the starting policy; a preference signal, a learned reward model, or a programmatic verifier then selects the optimization family.
Figure 1 shows the structure. A pretrained base model is fine-tuned on demonstrations. From there, pairwise preference data feeds the DPO family, a reward model feeds PPO-style RLHF, and a verifier feeds GRPO-style reinforcement learning with verifiable rewards (RLVR). All three converge on an aligned model that is then evaluated and red-teamed, and that evaluation loop frequently sends you back to collect more data.
In short: SFT teaches the format and baseline behavior from demonstrations, RLHF optimizes a learned reward with online reinforcement learning, DPO optimizes the same underlying preference objective offline with a classification-style loss, and RLVR replaces the learned reward with a program that checks the answer.
Notation used below
Let $x$ be a prompt, $y$ a response, $\pi_\theta$ the policy being trained, and $\pi_{ref}$ a frozen reference policy, usually the SFT model. A preference pair is $(x, y_w, y_l)$ where $y_w$ was preferred over $y_l$. A reward function is $r(x, y)$. The sigmoid is $\sigma$. Every loss below can be written in these terms, which makes the relationships between methods easy to see.
Method 1: Supervised Fine-Tuning
SFT is next-token cross-entropy on a curated dataset of prompt and response pairs. The loss is the negative log-likelihood of the demonstrated response, usually computed only on response tokens and masked on the prompt:
$$\mathcal{L}{SFT}(\theta) = -\,\mathbb{E})$$}} \sum_{t=1}^{|y|} \log \pi_\theta(y_t \mid x, y_{<t
There is no sampling from the model, no reward, and no reference model. A training step is a forward and backward pass over fixed data, which is why SFT is the cheapest and most predictable stage and the only one you can run comfortably with parameter-efficient adapters on a single GPU.
What SFT is good at
SFT is unmatched at teaching format and capability gaps: chat templates, tool-call syntax, JSON schemas, domain vocabulary, and house style. If the model has never produced the behavior, a demonstration is the most direct way to show it. Data quality dominates quantity here. Published practice and several open recipes emphasize curated, diverse instruction data over raw volume, and small high-quality sets can move a model substantially.
SFT is also the best defense against a subtle failure of the later stages. DPO and PPO both regularize toward a reference model. If that reference is a weak or off-distribution model, the regularizer pulls the policy toward something you did not want. A good SFT model makes everything downstream better behaved.
What SFT cannot do
SFT only ever raises the probability of demonstrated responses. It never says “this answer is worse than that one,” so it cannot teach a model to avoid a plausible-but-wrong behavior unless that behavior is absent from the data. Two consequences follow. First, SFT inherits the biases of its annotators or teacher model, such as a preference for long answers. Second, imitation encourages confident style even when the model lacks the knowledge, a known contributor to hallucination: the model learns to sound like an expert whether or not it can reproduce the expertise.
There is also an exposure problem. During training the model conditions on gold prefixes, but at inference it conditions on its own samples. Errors compound in ways SFT never penalizes. On-policy methods, which score the model’s own outputs, address precisely that gap, and it is a large part of why the field kept going past SFT.
SFT data practicalities
Typical SFT sets range from a few thousand to a few hundred thousand examples, depending on how broad the target behavior is. Mix in general instruction data when specializing, or the model will forget capabilities it was not shown (catastrophic forgetting). Train for a small number of epochs, because overfitting on a small set produces brittle, repetitive outputs. If compute is the constraint, adapter methods let you SFT a large model cheaply, and a single base can host many task adapters, a pattern we cover in multi-LoRA serving architecture.
RLHF: Reward Model Plus Online Reinforcement Learning
RLHF converts comparative human judgments into a scalar reward, then uses reinforcement learning to push the policy toward higher reward without drifting too far from the reference. The idea of learning rewards from human preference comparisons predates language models, appearing in deep reinforcement learning work on Atari and simulated robotics, and OpenAI and others carried it to text. The standard pipeline has three stages, shown in Figure 2.

Figure 2: The RLHF loop. Stages 1 to 4 build the reward model; stage 5 is an online loop in which the policy generates, the reward model scores, a KL penalty restrains drift, a critic estimates advantage, and a clipped update is applied.
The diagram makes the cost visible. Stages 1 to 4 are offline and relatively cheap. Stage 5 repeats thousands of times and, at each iteration, requires generation from the current policy, scoring by the reward model, a value estimate from a critic, and a gradient step. That is the part DPO removes.
Stage A: the reward model
Humans (or a stronger model) rank multiple candidate answers per prompt. A reward model, typically the SFT model with a scalar head, is trained on these comparisons using the Bradley-Terry model, which says the probability that $y_w$ beats $y_l$ is a sigmoid of the reward difference:
$$P(y_w \succ y_l \mid x) = \sigma\big(r_\phi(x,y_w) – r_\phi(x,y_l)\big)$$
$$\mathcal{L}{RM}(\phi) = -\,\mathbb{E} \log \sigma\big(r_\phi(x,y_w) – r_\phi(x,y_l)\big)$$
The reward model only needs to get the ordering right; absolute reward values are arbitrary up to a shift per prompt. That detail matters later, because it is exactly the invariance DPO exploits.
Stage B: KL-regularized policy optimization
The policy is then trained to maximize expected reward while staying close to the reference policy:
$$\max_{\pi_\theta}\; \mathbb{E}{x\sim\mathcal{D},\,y\sim\pi\theta(\cdot|x)}\big[r_\phi(x,y)\big] \;-\; \beta\, \mathrm{KL}\big[\pi_\theta(\cdot|x)\,|\,\pi_{ref}(\cdot|x)\big]$$
The coefficient $\beta$ controls how far the policy may move. Set it too low and the policy exploits the reward model, producing outputs that score highly for reasons that do not match human intent. This is reward hacking or over-optimization, and it is the central failure mode of RLHF. Set it too high and the policy barely changes from SFT.
The optimizer in the classic recipe is Proximal Policy Optimization. The PPO paper introduced a clipped surrogate objective that allows multiple epochs of minibatch updates per batch of samples while keeping each update inside an approximate trust region. In the language-model setting, the per-token reward is usually the KL penalty at every position, with the reward model’s score added at the final token, and a learned value function (the critic) estimates advantages, often via generalized advantage estimation.
Why RLHF is operationally heavy
A standard PPO RLHF step needs four models in memory: the trainable policy, the frozen reference, the reward model, and the critic. The critic is typically as large as the policy. Generation is the throughput bottleneck, since sampling long responses token by token is far slower than a forward pass over fixed data, and it usually needs a dedicated inference engine running alongside the trainer. PPO also has many interacting hyperparameters (learning rates for actor and critic, clip range, GAE parameters, KL coefficient, batch composition), and the results are sensitive to implementation details.
That sensitivity is real, but it should not be overstated. A 2024 study, Is DPO Superior to PPO for LLM Alignment?, argued that DPO has fundamental limitations and that, once PPO’s key implementation factors are tuned, PPO surpassed the other alignment methods in all of their dialogue and code experiments, including state-of-the-art results on challenging code competitions. The takeaway is not that PPO always wins. It is that many early “DPO beats PPO” comparisons measured an under-tuned PPO.
Cheaper online RL: REINFORCE-style estimators
A line of work questioned whether PPO’s machinery is needed at all for RLHF. The paper Back to Basics argued that many PPO components are unnecessary in this setting and that REINFORCE-style variants can outperform PPO and “RL-free” methods such as DPO and RAFT at lower cost. The practical result is the REINFORCE leave-one-out (RLOO) family: sample several responses per prompt, use the mean reward of the other samples as the baseline, and drop the critic entirely. That idea feeds directly into GRPO, covered below.
DPO vs RLHF: What Direct Preference Optimization Actually Does
The central insight of DPO is algebraic. The KL-regularized objective above has a known closed-form optimal policy, and you can invert that relationship to express the reward as a function of the policy itself.
For the objective in Stage B, the optimal policy satisfies:
$$\pi^*(y\mid x) = \frac{1}{Z(x)}\,\pi_{ref}(y\mid x)\,\exp!\Big(\tfrac{1}{\beta} r(x,y)\Big)$$
Rearranging gives the reward in terms of the policy:
$$r(x,y) = \beta \log \frac{\pi^*(y\mid x)}{\pi_{ref}(y\mid x)} + \beta \log Z(x)$$
Substitute this into the Bradley-Terry preference probability. The intractable partition function $Z(x)$ depends only on the prompt, so it cancels in the reward difference. What remains is a preference model written purely in terms of the policy and reference, and maximizing its likelihood gives the DPO loss:
$$\mathcal{L}{DPO}(\theta) = -\,\mathbb{E}\Big)$$} \log \sigma!\Big(\beta \log\frac{\pi_\theta(y_w|x)}{\pi_{ref}(y_w|x)} – \beta \log\frac{\pi_\theta(y_l|x)}{\pi_{ref}(y_l|x)
This is a binary classification loss over preference pairs. The quantity $\beta \log(\pi_\theta/\pi_{ref})$ is called the implicit reward. Training increases the implicit reward margin between chosen and rejected responses, and the reference model acts as an anchor, preventing the policy from drifting arbitrarily.
What you gain and lose
You gain a pipeline that looks like SFT. No sampling during training, no reward model, no critic, no rollout engine. You need the policy and a reference, and since the reference is frozen, you can precompute its log-probabilities once and then hold only the policy in memory during training. That is the origin of the often-quoted cost savings. The exact savings depend on model size, sequence length and hardware, so treat any specific percentage you read (including ones in older versions of this post) as setup-specific rather than universal.
You lose three things. First, DPO is offline: it learns from a fixed dataset generated by some other model, so the preference data’s distribution may differ from what your policy currently produces. Second, it has no explicit reward model to inspect, filter on, or reuse for best-of-n sampling and evaluation. Third, as the DPO-versus-PPO study above argues, its theoretical guarantees assume the reward is representable and the data covers the policy’s support, which is not true in practice.
The beta hyperparameter
In DPO, $\beta$ plays the same role as the KL coefficient in RLHF. Small $\beta$ lets the policy move far from the reference; large $\beta$ keeps it close. Common starting points in open implementations sit around 0.1, and sweeping a small range is standard practice, but the right value interacts with learning rate, data quality and the length of training. There is no universal setting, so tune it against held-out evaluations rather than the training loss.
Beyond DPO: IPO, KTO, ORPO and SimPO
DPO’s success produced a crowd of variants, each fixing a specific weakness. The useful way to read them is by which assumption they relax.
IPO: fixing overfitting to the preference model
The ΨPO paper from DeepMind identified two approximations behind RLHF and DPO: that pairwise preferences can be replaced by pointwise rewards, and that a reward model generalizes from collected data to the policy’s own samples. DPO removes the second but still leans on the first. When preferences in the data are nearly deterministic, Bradley-Terry pushes the reward gap toward infinity, and the KL anchor stops regularizing effectively. The paper’s Identity instantiation, widely called IPO, replaces the log-sigmoid with a squared loss that regresses the log-ratio margin toward a fixed target related to $1/(2\beta)$. The effect is a bounded objective that resists overfitting on clean, near-deterministic pairs. The paper reports empirical superiority to DPO on illustrative examples, so treat IPO as a robustness option rather than a proven general upgrade.
KTO: binary feedback instead of pairs
The KTO paper (ICML 2024) frames alignment through Kahneman and Tversky’s prospect theory. It shows that DPO and related losses are “human-aware losses” and proposes an objective that maximizes the utility of generations directly, requiring only a binary label per example (desirable or undesirable) rather than a ranked pair. The authors report matching or exceeding preference-based methods from 1B to 30B parameters. The practical consequence is data collection: thumbs-up and thumbs-down logs from a deployed product are far more abundant than matched pairs, and KTO can consume them directly.
ORPO: folding preference into SFT
ORPO argues that a small penalty on the disfavored generation style is enough when applied during SFT itself. It adds an odds-ratio term to the SFT loss, so one training stage replaces SFT followed by DPO and no reference model is needed. The authors report results on Phi-2, Llama-2 7B and Mistral 7B trained on UltraFeedback, including AlpacaEval 2.0 12.20%, IFEval 66.19% and MT-Bench 7.32 for their largest configuration. Those numbers are the authors’ own and are specific to one dataset and judge, so do not extrapolate them. The appeal is operational: fewer stages and lower memory.
SimPO: reference-free and length-normalized
SimPO (NeurIPS 2024) uses the average log-probability of a sequence as the implicit reward, which removes the reference model and aligns the training objective with how models are scored at generation time. It adds a target reward margin to the Bradley-Terry objective. The authors report gains over DPO of up to 6.4 points on AlpacaEval 2 and up to 7.5 points on Arena-Hard across Mistral, Llama 3 and Gemma 2 models. As with ORPO, these are self-reported results on specific benchmarks, and length-controlled judging is exactly where preference methods tend to disagree, so run your own evaluation before committing.
A pattern across the variants
Notice what these methods share. All are offline. All optimize a margin between chosen and rejected outputs. All differ in how they bound, normalize or anchor that margin. They are variations on one theme, which is why the choice among them matters less than the choice between offline preference optimization and online reinforcement learning, the subject of the next section.
Online RL for Reasoning: GRPO and RLVR
The 2025 shift was not a better preference loss. It was dropping the preference signal for domains where a program can check the answer. Mathematics has a ground-truth final answer. Code has unit tests. Many structured tasks have schema validators. For these, you can compute a reward without humans and without a learned reward model, which removes both the labeling cost and the reward-hacking surface of a neural judge.
GRPO: dropping the critic
Group Relative Policy Optimization was introduced in the DeepSeekMath paper as a PPO variant that improves mathematical reasoning while reducing PPO’s memory use. Instead of a learned value function, GRPO samples a group of $G$ responses for each prompt, scores each, and uses the group’s own statistics as the baseline. The advantage of response $i$ is its reward standardized within the group:
$$\hat{A}_i = \frac{r_i – \mathrm{mean}({r_1,\dots,r_G})}{\mathrm{std}({r_1,\dots,r_G})}$$
The policy then receives a PPO-style clipped update using this advantage, with a KL term toward the reference. Removing the critic eliminates a model as large as the policy, which is the headline efficiency win. DeepSeekMath 7B reached 51.7% on the MATH benchmark without tools, as reported in that paper.

Figure 3: Offline preference optimization trains on a fixed dataset with a classification-style loss. Online methods close a loop in which the current policy generates rollouts that a reward model or verifier scores.
Figure 3 shows the structural split that matters most. The offline branch never asks the current policy to generate anything during training. The online branch does, so its training data always reflects the policy’s current behavior. That is the property that lets online methods keep improving after offline datasets saturate.
RLVR and the DeepSeek-R1 result
Reinforcement learning with verifiable rewards uses a rule-based checker as the reward. The DeepSeek-R1 paper, later published in Nature, reported that reasoning abilities can be incentivized through pure reinforcement learning without human-labeled reasoning trajectories, with behaviors such as self-reflection and verification emerging during training. It reported superior performance on verifiable tasks such as mathematics, coding competitions and STEM fields relative to counterparts trained on human demonstrations, and that reasoning patterns of large models can be distilled into smaller ones.
Open recipes followed. The Tulu 3 paper released the data, code and recipe for a post-training pipeline built on Llama 3.1 that applies SFT, then DPO, then RLVR, and reports surpassing instruction-tuned Llama 3.1, Qwen 2.5 and Mistral as well as GPT-4o-mini and Claude 3.5 Haiku on its evaluation suite. It also documents training methods that did not reliably improve performance, which is rarer and more useful than another success story. The DAPO paper open-sourced a large-scale RL system reaching 50 points on AIME 2024 from the Qwen2.5-32B base model, and we cover another open reinforcement learning stack in our analysis of Xiaomi MiMo-V2.6’s open-weight RL training.
GRPO’s known biases
GRPO is not free of problems. The paper Understanding R1-Zero-Like Training identified an optimization bias that artificially increases response length, especially for incorrect outputs, and proposed Dr. GRPO to remove it. The same paper found that some base models (DeepSeek-V3-Base, Qwen2.5) already show substantial reasoning ability before any RL. That is an important caveat: part of what RL appears to “teach” may be elicitation of capability already acquired in pretraining. The paper also reports 43.3% on AIME 2024 with a minimalist 7B recipe.
Where RLVR does not apply
Verifiers only cover what can be verified. Tone, helpfulness in open-ended conversation, creative writing and nuanced safety judgments do not have a program that returns the right answer. For these, you are back to preference data, and either a reward model or a DPO-style loss. A related option is learning from AI feedback. The Constitutional AI paper used a list of written principles as the only human oversight: a model critiques and revises its own outputs for the supervised phase, and an AI model’s preference labels train a preference model for the RL phase, a technique the authors call RL from AI Feedback. It reduces human labeling needs for harmlessness, at the price of inheriting the labeling model’s biases.
Head-to-Head: Cost, Data, Stability and Control
With the mechanisms in place, we can compare the methods on the dimensions that decide real projects. The table below is qualitative on purpose. Published numbers depend on model size, dataset, judge and hyperparameters, and I would rather give you the structure than a benchmark table that looks precise and is not.
| Dimension | SFT | RLHF (reward model + PPO) | DPO and offline variants | GRPO / RLVR |
|---|---|---|---|---|
| Signal | Demonstrations | Ranked comparisons via a learned reward | Preference pairs (KTO: binary labels) | Programmatic verifier |
| Online sampling during training | No | Yes | No | Yes |
| Models in memory | Policy | Policy, reference, reward model, critic | Policy, reference (or precomputed log-probs) | Policy, reference, verifier code |
| Dominant failure mode | Imitates annotator bias, hallucinated confidence | Reward hacking, instability | Likelihood displacement, length bias, offline mismatch | Length bias, reward gaming, verifier gaps |
| Hyperparameter sensitivity | Low | High | Moderate | Moderate to high |
| Best fit | Format, capability, style | Open-ended quality where a good reward model exists | Fast preference tuning with existing data | Math, code, checkable tasks |
| Implementation effort | Low | High | Low to moderate | Moderate to high |
A worked memory estimate (illustrative arithmetic)
Consider a 7B-parameter model in bf16, where each full copy of the weights takes about 14 GB. Mixed-precision Adam training typically needs on the order of 16 bytes per trainable parameter once you count weights, gradients, and fp32 optimizer state, which is about 112 GB for one trainable 7B model before activations. This is a rule-of-thumb accounting, not a measurement of any specific framework.
A PPO RLHF setup trains two models (policy and critic) and holds two frozen ones (reference and reward model). That is roughly 224 GB of trainable state plus 28 GB of frozen weights, before activations and the generation engine’s KV cache. A DPO setup trains one model and holds one frozen reference, about 112 GB plus 14 GB, and if you precompute reference log-probabilities it drops to the trainable model alone. These figures explain why DPO fits on hardware where PPO needs sharding across many accelerators, and why adapter methods (LoRA) shrink the trainable state dramatically for both. The ratio matters more than the absolute numbers: PPO needs roughly twice DPO’s trainable state, plus a rollout engine.
Data requirements
SFT needs high-quality demonstrations, which are expensive per example because a human or a strong model must write a full answer. Preference data is cheaper per judgment, since choosing the better of two answers is easier than writing one, but you need many more of them, and the pairs should be on-policy or near-policy to be useful. KTO relaxes this further to unpaired binary labels. RLVR needs prompts with verifiable answers and a trustworthy checker, no human labels at all, but building a robust verifier is real engineering work.
Stability in practice
Offline methods are stable in the sense that training is a supervised-learning-style optimization over a fixed set: loss curves behave, runs are reproducible, and failures show up as degraded evaluations rather than divergence. Online RL can diverge, collapse in entropy, or exploit the reward in ways that only appear a few hundred steps in. That is the cost of its higher ceiling. If your team has not run RL before, budget extra time for monitoring: reward statistics, KL to reference, response-length distribution and entropy are the four curves to watch.
What Do Modern Labs Actually Use?
Disclosure is partial. Frontier labs rarely publish full post-training recipes, so what follows is limited to what papers state.
The Llama 3 paper describes a post-training approach built from supervised fine-tuning, rejection sampling and direct preference optimization, iterated over several rounds, rather than PPO. As I recall the paper, the authors reported choosing DPO over PPO for compute and stability reasons; I have flagged that detail as not independently re-verified in this rewrite. The Llama 3 paper itself describes a family whose largest model is a dense 405B Transformer with a 128K token context window.
Tulu 3 uses SFT, DPO and RLVR, as noted. DeepSeek-R1 emphasizes reinforcement learning against verifiable rewards for reasoning. InstructGPT used SFT plus PPO-based RLHF. Constitutional AI used AI-generated preferences. Taken together, the pattern across disclosed recipes is: SFT first, always; an offline preference stage for conversational quality and safety; and an online RL stage where verifiable rewards exist. Treat anything stronger as speculation about closed systems. Our post on OpenAI o3 and o4-mini reasoning models discusses what is and is not public about reasoning-model training.
The sequencing logic
Why this order? SFT puts the model in a region where its samples are decent, which is a precondition for both preference learning (the reference must be sensible) and RL (sparse rewards are unlearnable if the policy never succeeds). The preference stage then sharpens style and refusal behavior cheaply. The RL stage then pushes capability on tasks you can check. Each stage consumes a different signal, which is why the methods stack more often than they compete.
A Practitioner Decision Procedure

Figure 4: Choose by signal type first, cost tolerance second. SFT is always the starting point; the rest depends on whether you hold pairs, binary labels, or a verifier.
Figure 4 encodes the procedure. Start from SFT unless your base model already follows instructions well. If your only data is demonstrations, stop there and evaluate. If you have preference pairs and want low cost and low risk, use DPO, and consider IPO or SimPO if you see overfitting or length drift. If your feedback is thumbs up or down, use KTO. If you have a checker, use GRPO-style RL. Reach for a reward model plus online RL only when you need to optimize a quality signal that is open-ended, you have the engineering capacity for RL, and you have evidence that offline methods have plateaued.
Evaluating alignment honestly
Whatever you pick, the evaluation protocol decides whether you can trust the result. Hold out prompts that were never used for training or for hyperparameter selection. Measure length alongside quality, because many judge-based benchmarks reward longer answers, and use length-controlled comparisons where available. Check safety regressions explicitly, since optimizing for helpfulness can erode refusals. For agentic systems, trajectory-level evaluation matters more than single-turn scoring, which we discuss in our piece on AI agent evaluation harnesses. Finally, do not trust a single judge model: calibrate it against human spot checks.
Trade-offs, Gotchas, and What Goes Wrong
Likelihood displacement in DPO. DPO is meant to raise the probability of chosen responses, but the loss only constrains the margin, so both probabilities can fall. The paper Unintentional Unalignment documents that the likelihood of preferred responses often decreases during training and that probability mass can shift toward unintended outputs. In one safety experiment it reports refusal rates for Llama-3-8B-Instruct dropping from 74.4% to 33.4% after DPO on refusal preferences. The cause they identify is high embedding similarity between chosen and rejected responses, and filtering such pairs mitigated the problem. Practical lesson: audit your pairs, and monitor chosen-response log-probability, not just the margin.
Length and style exploitation. Both reward models and DPO margins correlate with length in many datasets. A policy can win by being verbose. Length normalization (SimPO), length-controlled evaluation, and explicit length penalties are the usual defenses.
Reward hacking. Any learned reward is a proxy. Push hard enough and the policy finds inputs where the proxy and the truth diverge. Mitigations include a stronger KL penalty, early stopping on a held-out gold signal, reward model ensembles, and refreshing the reward model on current policy samples.
Offline-online mismatch. If the preference pairs came from a different model than the one you train, DPO is learning about outputs your policy rarely produces. Regenerating pairs with the current policy, scoring them, and training again (iterative or online DPO) is a common fix, and it is also a step toward the online regime.
Verifier gaming. RLVR policies exploit weak checkers: emitting the answer format without reasoning, hard-coding test cases, or probing edge cases in a sandbox. Hold-out tests and format-independent verification reduce this. For code execution rewards, run untrusted code in isolation, a topic we cover in our comparison of agent sandboxes.
Reproducibility. Results in this literature are famously sensitive to base model, data mix, judge and seed. A method that wins on one benchmark configuration can lose on another, and the DPO-versus-PPO debate is the clearest example. Plan your own ablation rather than importing someone else’s ranking.
Distillation and data provenance. Preference and SFT data generated by stronger models can carry licensing restrictions and leak behaviors from the teacher. Check the terms of any model whose outputs you train on.
Practical Recommendations
For a team doing its first alignment project, the lowest-risk path is SFT followed by DPO, with adapters if hardware is tight. That stack has well-supported open implementations, behaves predictably, and gives a measurable improvement on conversational quality and refusal behavior. Add online RL only after you have a trustworthy evaluation and evidence that you are leaving quality on the table.
If your product is about correctness on checkable tasks (code generation, structured extraction, math, data transformation), invest in the verifier before the algorithm. A good verifier plus a simple GRPO implementation will outperform a sophisticated algorithm against a weak checker. If your feedback arrives as thumbs up or thumbs down, KTO lets you use it without constructing pairs.
Keep the engineering honest: version datasets, log the four RL health curves, and keep a frozen evaluation set that never touches training. Checklist:
- Run SFT first and evaluate it as your baseline before any preference stage.
- Inspect preference pairs for near-duplicates, length skew and high similarity between chosen and rejected.
- Track chosen-response log-probability, KL to reference, response length and entropy.
- Sweep $\beta$ over a small range and select on held-out evaluation, not training loss.
- Evaluate with length-controlled metrics and an explicit safety regression suite.
- Treat every published benchmark delta as a hypothesis to reproduce on your own data.
- Use an isolated sandbox for any code-execution reward.
Frequently Asked Questions
Is DPO better than RLHF?
Neither dominates. DPO is simpler, cheaper and more stable, and the original paper reported it matching or exceeding PPO-based RLHF on sentiment control, summarization and dialogue. Later work showed a well-tuned PPO can beat DPO on dialogue and code tasks. DPO wins on cost and simplicity for offline preference data; online RLHF can win on peak quality when you can afford the engineering.
Do I need SFT before DPO or RLHF?
In practice, yes. DPO and RLHF both regularize toward a reference policy, and RL needs a policy that already produces reasonable outputs. SFT creates that starting point and teaches the format. Skipping it is possible when you start from an already instruction-tuned model, or with ORPO, which folds the preference signal into the SFT stage so a single stage replaces both.
What is the difference between PPO and GRPO?
Both use a clipped policy-gradient update. PPO learns a value function (critic) to estimate advantages, which is a second model as large as the policy. GRPO removes the critic and standardizes each response’s reward against the other responses sampled for the same prompt. That saves memory and works well with verifiable rewards, but it has known length biases that Dr. GRPO corrects.
How much preference data does DPO need?
There is no universal number. Published recipes range from tens of thousands of pairs for open chat models to far more for broad-coverage systems, and quality matters more than count. Pairs where chosen and rejected responses are very similar can cause likelihood displacement. Start with a clean, deduplicated set, evaluate, and scale only if held-out results keep improving.
Can I do alignment with LoRA instead of full fine-tuning?
Yes. SFT and DPO both work with low-rank adapters, which cut trainable state and memory sharply, and with a frozen base you can often use the adapter-disabled model as the reference without loading a second copy. Quality can trail full fine-tuning on some tasks, so compare on your evaluation set. Our guide to LoRA, QLoRA and full fine-tuning covers the trade-offs.
What is RLVR and when should I use it?
RLVR, reinforcement learning with verifiable rewards, uses a program that checks the answer, such as an exact-match math grader or unit tests, as the reward instead of a learned reward model. Use it for tasks with objectively checkable outcomes. It is less suitable for open-ended qualities like tone or helpfulness, which still need preference data.
Further Reading
Internal:
- LoRA vs QLoRA vs full fine-tuning vs distillation for small language models for the parameter-efficiency layer under every method above.
- Xiaomi MiMo-V2.6 open-weight RL stack explained for a recent open reinforcement learning recipe.
- OpenAI o3 and o4-mini reasoning models explained for what is public about reasoning-model training.
- AI agent evaluation harness and trajectory evals for evaluating aligned agents beyond single-turn scores.
External primary sources:
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (Rafailov et al.)
- Training language models to follow instructions with human feedback (Ouyang et al.)
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training (Allen Institute for AI)
By Riju — about

Pingback: GRPO and RLVR Explained: How Reasoning Models Are Trained wi