GRPO Reinforcement Learning and RLVR: How Reasoning Models Are Actually Trained
Most explanations of reasoning models stop at “they were trained with reinforcement learning.” That sentence hides the interesting part: a single design decision, throwing away the learned value network and using the other samples of the same prompt as the baseline, is what made the recipe cheap enough for open labs to reproduce. GRPO reinforcement learning is that decision. Paired with a reward that a program can check, known as RLVR (reinforcement learning with verifiable rewards), it turns a pretrained base model into one that writes long chains of thought, verifies its own work and backtracks.
This matters in 2026 because the same loop now sits under nearly every open reasoning model, and its failure modes are well documented: length inflation, entropy collapse, reward hacking and a base-model dependence that the headline results obscure. If you are going to run it, or just judge claims made about it, you need the mechanics rather than the slogan.
This post derives the GRPO objective from PPO, walks one group of rollouts with real numbers, shows a runnable TRL training script, and catalogues what goes wrong in practice.
What this covers: PPO versus GRPO, the group-relative advantage formula, the KL term and clipping, RLVR reward design for math and code, the DAPO and Dr. GRPO fixes, a minimal GRPOTrainer example, compute cost arithmetic, and the failure modes that cost real training runs.
Context and Background
Reinforcement learning from human feedback (RLHF) was the first widely deployed way to apply RL to language models. A reward model, trained on human preference pairs, scores responses, and Proximal Policy Optimization (PPO) updates the policy to raise that score while staying near the starting model. If you want the full comparison against direct preference methods, our DPO vs RLHF vs SFT alignment benchmark covers that ground.
Preference-based RLHF has two properties that make it a poor fit for reasoning. First, the reward is a learned proxy for human taste, so a policy optimized hard against it finds the proxy’s holes. Second, a human grader cannot cheaply tell whether a 6,000-token derivation is correct. For tasks with a ground truth, such as a numeric answer or a unit-test suite, a program is a better judge than a person or a neural reward model.
That observation is the core of RLVR. The reward is computed by a deterministic checker: does the final answer match, do the tests pass, does the output follow the required format. The term was popularized by the Tulu 3 post-training work from the Allen Institute for AI, and it became the default framing once DeepSeek published its results.
The algorithmic half arrived earlier. The DeepSeekMath paper (Shao et al., arXiv 2402.03300) introduced Group Relative Policy Optimization as a variant of PPO that improves mathematical reasoning while reducing PPO’s memory use. The paper’s headline model, DeepSeekMath 7B, reaches 51.7% on the MATH benchmark without external tools or voting, and 60.9% with self-consistency over 64 samples, according to the abstract.
The DeepSeek-R1 paper (arXiv 2501.12948, later published in Nature under DOI 10.1038/s41586-025-09422-z) then argued that RL alone, with no human-labeled reasoning demonstrations, can elicit behaviors such as self-verification and mid-solution strategy changes. Its R1-Zero experiment applied GRPO directly to a base model with rule-based rewards. That is the experiment everyone reproduced, and it is where our story starts.
Reinforcement learning is also a general control technology, not only a language-model trick. For a contrasting physical-world application of the same policy-gradient family, see our piece on reinforcement learning for tokamak plasma control, and for simulation-based policy training see the Isaac Lab robot training tutorial.
From PPO to GRPO: The Core Idea
Direct answer: GRPO is PPO without the critic. For each prompt it samples a group of G completions, scores each with the reward function, and uses the group’s own mean and standard deviation as the baseline. Each completion’s advantage is its reward minus the group mean, divided by the group standard deviation. No value network is trained.

Figure 1: The GRPO loop. A frozen reference model anchors the KL term, a verifier scores each of G sampled completions, and group statistics replace the learned value function.
The diagram shows the single structural change. Everything else, sampling from the current policy, a clipped surrogate objective, an optional KL anchor to a reference model, is inherited from PPO. The group baseline is what removes the largest component of the training system.
What PPO needs and why it is expensive
PPO estimates, for every token position, how much better the sampled token was than the policy’s average. That requires a value function V(s) that predicts expected future reward from any prefix. In LLM practice the value function is a second transformer, typically initialized from the policy or a reward model and of comparable size. Generalized Advantage Estimation then combines value predictions across time steps to produce a per-token advantage.
The costs are concrete. As illustrative arithmetic, mixed-precision Adam training costs roughly 16 bytes per trainable parameter (bf16 weights and gradients, fp32 master weights, and two fp32 moment buffers). For a 7B-parameter policy that is about 112 GB of optimizer-resident state, and a same-sized critic adds another 112 GB, before activations, a frozen reference model and the rollout engine’s KV cache. Dropping the critic halves the trainable state.
There is a second, subtler cost. The critic must learn something useful from a reward that arrives only once, at the end of a long sequence. With sparse terminal rewards, value estimates for early tokens are noisy, and a poorly trained critic injects bias into every policy update. GRPO sidesteps that by never asking the question.
The group-relative advantage
For a prompt q, sample G completions o_1 … o_G from the old policy and compute scalar rewards r_1 … r_G. The advantage assigned to every token of completion i is:
A_i = (r_i – mean(r_1..r_G)) / std(r_1..r_G)
This is a z-score within the group. A completion that beats its siblings gets a positive advantage, one that lags gets a negative one, and the magnitude is normalized so that easy and hard prompts produce comparably scaled gradients. All tokens in a completion share the same advantage, which is a deliberate simplification: credit assignment is coarse, and the method relies on volume of samples rather than token-level insight.
The baseline logic is the classical variance-reduction argument from policy gradients. Subtracting any baseline that does not depend on the sampled action leaves the gradient unbiased in expectation. The group mean is a Monte Carlo estimate of the prompt’s value under the current policy, computed from siblings rather than predicted by a network.
The full objective
The GRPO objective maximizes, averaged over prompts and the G samples, a PPO-style clipped term minus a KL penalty:
J = E[ (1/G) * sum_i (1/|o_i|) * sum_t ( min( rho_it * A_i, clip(rho_it, 1-eps, 1+eps) * A_i ) – beta * KL(pi_theta || pi_ref) ) ]
Here rho_it = pi_theta(o_it | q, o_i,<t) / pi_old(o_it | q, o_i,<t) is the importance ratio between the policy being updated and the policy that generated the samples. Clipping at 1 plus or minus eps prevents a single large step from moving the policy far from the sampler, which matters because the samples are only valid for a policy close to the one that produced them.
The KL term uses an unbiased per-token estimator: pi_ref/pi_theta – log(pi_ref/pi_theta) – 1. It is always non-negative and has low variance. Unlike classic RLHF, which folds the KL penalty into the reward, GRPO places it directly in the loss, so the group advantage stays a pure function of the verifier’s output.
One practical point deserves emphasis. In the current TRL documentation, beta defaults to 0.0, meaning the KL term is switched off and no reference model is loaded. Many recent recipes drop the KL anchor entirely for verifiable-reward training, because the reward cannot be “hacked” in the preference-model sense, and a reference model costs memory and a forward pass per step. Whether that is wise depends on your verifier, which we return to in the failure-modes section.
Deeper Analysis: A Worked Group, RLVR Rewards and a Runnable Trainer
The formulas are short; the behavior is in the details. The next sections follow a single group through the computation, then show how reward functions are written and how to launch a real run.
One group, with numbers
Take one prompt, G = 8 samples, and a binary verifier. Three completions reach the correct answer.
| Sample | Reward r | Advantage (population std) |
|---|---|---|
| 1 | 1 | +1.291 |
| 2 | 0 | -0.775 |
| 3 | 0 | -0.775 |
| 4 | 1 | +1.291 |
| 5 | 0 | -0.775 |
| 6 | 0 | -0.775 |
| 7 | 0 | -0.775 |
| 8 | 1 | +1.291 |
The group mean is 3/8 = 0.375. The population standard deviation is sqrt(0.375 * 0.625) = 0.484. A correct sample gets (1 – 0.375)/0.484 = +1.291; an incorrect one gets (0 – 0.375)/0.484 = -0.775. Implementations that use the unbiased sample standard deviation (dividing by G-1) get a slightly larger denominator of about 0.517 and slightly smaller advantages, roughly +1.21 and -0.73. The difference is cosmetic, but it explains why two implementations report different loss curves for identical data.
Now consider the two degenerate cases. If all eight samples are wrong, every reward is 0, the standard deviation is 0, and the advantage is 0/0, which implementations guard with a small epsilon so every advantage is effectively zero. If all eight are right, the same thing happens. In both cases the prompt contributes no gradient at all. This is not a corner case: as a model improves, a growing fraction of the training set becomes all-correct, and the effective batch shrinks. The DAPO paper’s dynamic sampling exists to counter exactly this.

Figure 2: An RLVR reward pipeline. Each completion passes through format, answer and execution verifiers, and the per-verifier scores are combined into one scalar.
Designing verifiable rewards
A verifiable reward is any function from completion (and a hidden ground truth) to a number that code can compute. The three families used in practice are answer matching, execution, and format compliance.
Answer matching is the math case. The model is told to put its final answer in a fixed place, commonly a LaTeX boxed expression, and the verifier extracts it and compares it to the reference. Naive string equality fails on “0.5” versus “1/2”, so serious verifiers normalize expressions or use a symbolic library. The R1 paper describes rule-based accuracy rewards of this kind, and deliberately avoided neural reward models for reasoning because of reward-hacking risk at scale.
Execution is the code case. The completion is a program, the verifier runs it against hidden unit tests in a sandbox, and the reward is the pass rate or a pass/fail bit. This is stronger than string matching but introduces infrastructure: sandboxed execution, timeouts, resource limits and protection against a model that learns to read the test file. Treat the sandbox as a security boundary, not a convenience.
Format rewards are small bonuses for obeying a template, such as placing reasoning inside think tags and the answer inside answer tags. They are cheap to compute and they stabilize early training, when the base model has not yet learned the output structure. They should stay a minor share of the total, otherwise the policy optimizes the template instead of the answer.
A minimal answer-and-format reward in the shape TRL expects looks like this. Reward functions receive the completions plus any extra dataset columns as keyword arguments and return one float per completion.
import re
BOXED = re.compile(r"\\boxed\{(.*?)\}", re.S)
THINK = re.compile(r"^<think>.*?</think>\s*.+$", re.S)
def normalize(s: str) -> str:
return s.strip().replace(" ", "").rstrip(".")
def accuracy_reward(completions, ground_truth, **kwargs):
"""1.0 if the boxed answer equals the reference, else 0.0."""
out = []
for text, gt in zip(completions, ground_truth):
m = BOXED.search(text)
out.append(1.0 if m and normalize(m.group(1)) == normalize(gt) else 0.0)
return out
def format_reward(completions, **kwargs):
"""Small bonus for a think block followed by an answer."""
return [0.2 if THINK.match(c) else 0.0 for c in completions]
Note that when the dataset uses a conversational prompt format, TRL passes each completion as a list of message dictionaries rather than a plain string, so check the structure of your inputs before relying on regular expressions. The sketch above assumes plain-text completions.
A runnable TRL GRPOTrainer example
The Hugging Face TRL library ships a GRPOTrainer. Its documentation’s minimal example trains Qwen/Qwen2.5-0.5B-Instruct on the trl-lib/DeepMath-103K dataset with the built-in accuracy_reward, and notes that distributed training across 8 GPUs takes about a day. The script below extends that pattern with a custom reward list and explicit configuration. Parameter names follow the TRL documentation at the time of writing; check the version you install, because this API has changed quickly.
# train_grpo.py -- run with: accelerate launch train_grpo.py
from datasets import load_dataset
from trl import GRPOConfig, GRPOTrainer
from trl.rewards import accuracy_reward
dataset = load_dataset("trl-lib/DeepMath-103K", split="train")
args = GRPOConfig(
output_dir="qwen-grpo",
num_generations=8, # G: completions per prompt
max_completion_length=1024, # cap on reasoning length
beta=0.0, # KL weight; 0.0 = no reference model
loss_type="dapo", # token-level normalization
scale_rewards=True, # divide advantages by group std
learning_rate=1e-6,
per_device_train_batch_size=8,
logging_steps=10,
)
trainer = GRPOTrainer(
model="Qwen/Qwen2.5-0.5B-Instruct",
reward_funcs=[accuracy_reward, format_reward], # summed by default
args=args,
train_dataset=dataset,
)
trainer.train()
Several parameters deserve a comment. num_generations is G; larger groups give a better baseline estimate and a higher chance that a hard prompt contains at least one success, at linear rollout cost. loss_type selects the normalization scheme, discussed next. scale_rewards controls whether advantages are divided by the group standard deviation, and TRL also offers a batch-level option. With multiple reward functions, rewards are summed unless you set reward_weights.
The loss normalization question: GRPO, Dr. GRPO and DAPO
The original objective divides each completion’s token sum by its own length, 1/|o_i|. That looks harmless but creates a length bias. Because a long wrong answer is divided by a large number, each of its tokens is penalized less than the tokens of a short wrong answer, so the policy is nudged toward verbose failures.
The Dr. GRPO paper (arXiv 2503.20783) identifies this optimization bias, reporting that GRPO inflates response length during training, especially for incorrect answers. It proposes a corrected objective that, per the TRL documentation, divides by a constant such as the maximum completion length instead of the sequence length. The same paper also reports a pretraining bias: the DeepSeek-V3-Base model already exhibits an “aha moment”, and Qwen2.5 base models reason well even without prompt templates. Its minimalist R1-Zero recipe reaches 43.3% on AIME 2024 with a 7B base model, which the authors present as state of the art for that size.
The DAPO paper (arXiv 2503.14476) reports 50 points on AIME 2024 with a Qwen2.5-32B base model and introduces four techniques. From the paper, these are Clip-Higher (a decoupled, asymmetric clipping range that lets low-probability tokens grow), Dynamic Sampling (discarding groups with all-equal rewards), a token-level policy gradient loss, and Overlong Reward Shaping (soft penalties for truncated outputs). TRL exposes the token-level scheme as loss_type="dapo" and the Dr. GRPO scheme as loss_type="dr_grpo".

Figure 3: How the three variants differ. They share the group advantage and clipped surrogate but change normalization, clipping range and which groups are kept.
How DeepSeek-R1 used the pieces
The R1 paper describes two lines of work. R1-Zero applies RL directly to the base model with rule-based rewards and no supervised warm-up. It learns to produce longer chains of thought and to re-check itself, but its outputs are hard to read and mix languages, which is why the released R1 model uses a multi-stage pipeline: a small cold-start dataset of long reasoning examples, reasoning-focused RL, rejection sampling to build a new supervised set, and a final RL stage that adds helpfulness and safety objectives.
The first version of the paper reported that R1-Zero’s pass@1 on AIME 2024 climbed from 15.6% to 71.0% over training, and 86.7% with majority voting. Those figures come from my reading of the original January 2025 version; the paper was revised in January 2026 for its journal publication, so verify them against the current text before citing. The qualitative pattern is the durable point: accuracy rises together with response length, because longer deliberation is rewarded whenever it leads to a verified answer.
The paper also reports that smaller models improve more by distillation from R1’s reasoning traces than by running the same RL on the small model directly. For practitioners that is a significant budget signal. RL is the expensive step to do once at scale, and supervised fine-tuning on its outputs is the cheap way to spread the result.
If you want to see how a recent frontier reasoning model packages these ideas, our Kimi K3 architecture and benchmark breakdown is a useful companion.
Compute, Sequence Diagram and Cost
The mechanics are cheap to state and expensive to run. The cost structure is dominated by generation, not by gradient computation, and that changes how you provision hardware.

Figure 4: One training step as a sequence. Rollout engines generate G completions per prompt, verifiers score them, and the learner updates the policy and syncs weights back.
Where the tokens go
Consider an illustrative step: 256 prompts, G = 16 completions each, and an average completion length of 4,000 tokens. That is 256 x 16 x 4,000 = 16.4 million generated tokens per step. A forward-and-backward pass over the same tokens costs roughly three times the forward-only compute, but autoregressive generation is memory-bandwidth bound and far less efficient per token than a batched training pass. In practice, rollout generation often takes the majority of wall-clock time per step, which is why serious setups run a dedicated inference engine such as vLLM alongside the trainer.
Three consequences follow. First, the long tail of completion lengths wastes capacity: a batch waits for its slowest sample unless the system uses asynchronous or partial-rollout scheduling. Second, weights must be synchronized from the trainer to the inference engine each step, a real cost for large models. Third, the memory saved by removing the critic is partly consumed by the KV cache for long contexts, so the net saving depends on sequence length.
What the group size buys
Group size trades signal quality against rollout cost. If a prompt has per-sample success probability p, the chance that a group of G has mixed outcomes, and therefore a nonzero gradient, is 1 – p^G – (1-p)^G. For p = 0.1, G = 4 gives about 34% and G = 16 gives about 81%. For p = 0.5, even G = 4 gives 88%. Hard prompts, where learning is most valuable, are exactly where a small group fails to produce any signal.
This arithmetic explains why curriculum matters. Filtering your dataset to prompts the current model solves between roughly 10% and 90% of the time keeps most groups informative. Dynamic sampling automates this by resampling until the batch is full of mixed groups, at the price of extra generation.
Reported training cost
The DeepSeek-R1 Nature publication reportedly states a compute cost of roughly 294,000 US dollars for the R1 reinforcement learning stage, on top of the base model’s own pretraining. I have not independently verified that figure against the published supplementary material, and it excludes the base model, data work and ablation runs, so treat it as a lower bound on what a first-time team will spend. Open reproductions on 7B-scale models report far smaller budgets, but benchmark numbers on small models should not be extrapolated to frontier scale.
Trade-offs, Gotchas, and What Goes Wrong
GRPO reinforcement learning is simple to implement and easy to run incorrectly. The following failure modes account for most of the lost training runs reported in public write-ups and issue trackers.
Length inflation. With the original per-sequence normalization, wrong answers get longer over time because long failures are penalized less per token. The result looks like “the model is thinking harder” on a response-length chart while accuracy is flat. Use a token-level or constant normalization, and track accuracy per token budget instead of response length alone.
Entropy collapse. The policy’s output distribution narrows quickly, exploration stops, and the group sees identical samples with identical rewards, which yields zero advantage. DAPO’s Clip-Higher addresses this by loosening the upper clipping bound so rare but useful tokens can gain probability. Monitor policy entropy and the fraction of groups with zero reward variance; if either trends the wrong way, you are training on a shrinking batch.
Reward hacking. Verifiable does not mean unhackable. A math checker that compares the first number in the output rewards a model that lists many numbers. A code checker that runs visible tests rewards special-casing. Hidden tests, held-out verifiers, strict parsing and a periodic human audit of high-reward samples are the standard defenses.
Base-model dependence. The Dr. GRPO analysis shows that some base models already reason before any RL. A striking “aha moment” curve on one model family may reflect pretraining data rather than RL discovering a skill. Always compare against the base model with a good prompt, and test on a second family before generalizing a recipe.
Sparse reward on hard problems. If the model never solves a prompt, the group baseline provides nothing. RLVR can sharpen what the model can already sometimes do; it is much weaker at teaching what it cannot reach by sampling. For genuinely new capability, supervised data or distillation is still required.
Verifier-bound scope. Everything outside what you can check automatically, such as writing quality, safety nuance and open-ended advice, gets no signal from RLVR. Production pipelines therefore combine verifiable rewards with preference-based stages, which reintroduces the reward-model problems RLVR was meant to avoid.
Off-policy drift. Running several optimization epochs per generation batch, or generating with a slightly stale policy for throughput, makes the importance ratios unreliable. Keep the clipped objective, keep num_iterations low (the TRL default is 1), and log the clipped fraction.
Practical Recommendations
Start small and instrument heavily. A 0.5B to 7B instruct model, a dataset with a reliable checker and a few hundred steps will show you whether the loop works before you commit serious compute.
Choose the loss normalization deliberately. For long chain-of-thought training, prefer a token-level (dapo) or constant-normalized (dr_grpo) loss over the original per-sequence form. Leave beta at 0 unless you see the policy drifting into degenerate outputs, in which case a small positive value restores a reference anchor at the cost of memory.
Treat the data distribution as a hyperparameter. Pre-filter prompts by the current model’s pass rate, keep groups large enough to be informative on your hardest slice, and hold out an evaluation set that the training verifier never touched.
Budget for generation. Size your inference capacity first, then your learner. Decide early whether to run synchronous or asynchronous rollouts, because it shapes the rest of the system.
A short launch checklist:
- Verifier tested on known-good and known-bad samples, including adversarial ones.
- Reward range small and bounded; format bonus no more than a fifth of the accuracy reward.
- Dashboards for mean reward, reward standard deviation, zero-variance group fraction, entropy, mean and 95th percentile completion length, clipped ratio.
- Held-out benchmark evaluated every N steps with a fixed decoding configuration.
- Baseline comparison: base model with best-effort prompting, and a supervised fine-tuning control.
- Checkpoints saved frequently, because collapse is often sudden.
Frequently Asked Questions
What is GRPO in reinforcement learning?
Group Relative Policy Optimization is a policy-gradient algorithm derived from PPO. For each prompt it samples several completions, scores them, and computes each completion’s advantage as its reward minus the group mean, divided by the group standard deviation. Because the baseline comes from the group, no separate value network is trained, which cuts memory and removes a source of bias. It was introduced in the DeepSeekMath paper and popularized by DeepSeek-R1.
What is the difference between GRPO and PPO?
PPO trains a value network to estimate per-token advantages through Generalized Advantage Estimation. GRPO replaces that critic with a Monte Carlo baseline computed from several samples of the same prompt and gives every token in a completion the same advantage. Both use a clipped importance ratio for stable updates. GRPO is cheaper in memory and simpler to run, but gives coarser credit assignment and needs enough samples per prompt for a reliable baseline.
What does RLVR mean?
RLVR stands for reinforcement learning with verifiable rewards. Instead of a learned reward model approximating human preference, the reward comes from a deterministic checker: an answer comparator for math, a unit-test runner for code, or a format validator. It is only applicable where correctness can be checked automatically, and it is far harder to exploit than a neural reward model, though a sloppy checker can still be gamed.
Do reasoning models need human-labeled reasoning data?
Not necessarily. The DeepSeek-R1 paper argues that RL with rule-based rewards can elicit reasoning behavior without human reasoning demonstrations, and R1-Zero is the evidence. In practice the released model still used a small cold-start dataset and a multi-stage pipeline for readability. Later analysis suggests the outcome depends heavily on the base model, since some already show reasoning behavior before RL begins.
Why does GRPO make responses longer?
Two effects combine. Rewarded long-form reasoning genuinely improves accuracy on hard problems, so some growth is real. Separately, the original per-sequence length normalization penalizes tokens in long wrong answers less than in short ones, inflating length for failures. Dr. GRPO and DAPO’s token-level loss remove that second effect, so measure accuracy gains, not length alone.
Can I run GRPO on a single GPU?
Yes, for small models. The TRL GRPOTrainer can train sub-billion-parameter models on one GPU, especially with beta at 0 so no reference model is loaded, parameter-efficient adapters and short completion limits. Expect slow progress, since generation dominates. Meaningful reasoning gains on 7B-class models generally need multiple GPUs and an inference engine such as vLLM for rollouts.
Further Reading
Internal:
- DPO vs RLHF vs SFT alignment benchmark for the preference-based alternatives to verifiable rewards.
- Kimi K3 explained: reasoning model architecture and benchmarks for a recent reasoning model in context.
- Reinforcement learning for tokamak plasma control for policy gradients in a physical control system.
- Isaac Lab robot training tutorial for simulation-based RL practice.
External primary sources:
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (arXiv 2402.03300)
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv 2501.12948)
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale (arXiv 2503.14476)
- Understanding R1-Zero-Like Training: A Critical Perspective, Dr. GRPO (arXiv 2503.20783)
- Hugging Face TRL GRPOTrainer documentation
By Riju — about
