LLMs
AI Evaluation

GRPO Explained: Group Relative Policy Optimization and the Data It Consumes

GRPO is the reinforcement learning algorithm behind DeepSeek-R1, Qwen3 and DeepSeek-V3.2. This guide explains how group relative policy optimization scores multiple responses against each other, and why your prompts and your reward model decide whether it learns.

Table of contents

AI Summary

  • GRPO drops PPO's critic and uses the group's mean score as the baseline.
  • If every sampled answer gets the same score, that prompt produces zero learning signal.
  • Qwen3's reasoning RL used 3,995 selected query–verifier pairs.
  • Scorer errors land directly in the advantage, because GRPO builds its baseline from the scores themselves.

Introduction

GRPO (Group Relative Policy Optimization) is usually explained as a cheaper version of PPO (Proximal Policy Optimization). PPO is the algorithm behind the original reinforcement learning from human feedback (RLHF) pipelines, which train a reward model on human preference data. That is accurate and incomplete. GRPO removes one model from the training loop: the critic. In its place it puts a statistic, the average score of several answers the model gave to the same prompt. The maths gets simpler. The data gets more important.

GRPO does not remove the separate reward model. In the paper that introduced it, a learned reward model still scored the answers. DeepSeek-R1-Zero later removed the learned reward by switching to rule-based checks. That was a separate decision about the data, not a property of the algorithm.

Separating the two removals shows where GRPO's training signal comes from. Without a critic, it depends on two inputs: which prompts go into each group, and what scores the group. Both are data-design questions.

The sections below cover the mechanism in one screen, then how labs have used GRPO at scale and the fixes production versions carry. Most of the article is about those two inputs.

What Is GRPO and What Does It Remove From PPO?

The DeepSeekMath paper (Shao et al., 2024) introduced GRPO as "a variant of Proximal Policy Optimization (PPO)." DeepSeekMath is a model specialised in mathematics, so GRPO started life as a way to train one domain-specific model. Its goals were stronger mathematical reasoning and lower memory usage than PPO.

To see what changed, start with PPO. PPO updates the policy model (the language model being trained) using an advantage: how much better an answer turned out than expected. "Expected" comes from a value function, a separate critic model trained alongside the policy. The critic estimates the cumulative reward still to come from each token, and PPO's advantage estimation combines those estimates using a discount factor for future rewards. The DeepSeekMath authors point out that this value function is "typically another model of comparable size as the policy model." You are training two large models to get one.

That is where GRPO's lower memory overhead comes from. Only the policy and the critic receive gradient updates in PPO, so removing the critic roughly halves the memory spent on models being trained. The total saving is smaller, because the reference model and any reward model still sit in memory.

GRPO drops the critic. For each prompt, GRPO samples a group of G responses to the same prompt from the current policy and gives each a reward score. The grouped rewards become the baseline:

A_i = (r_i − mean(r_1, …, r_G)) / std(r_1, …, r_G)

Each answer's advantage is its reward minus the group's mean reward, divided by the group's standard deviation. An answer is good if it beat its siblings, and the policy learns to maximize rewards relative to its own group. That comparative nature gives GRPO its name. Some baseline is always needed, because raw policy-gradient updates have high variance. What GRPO changes is where the baseline comes from.

The policy updates then follow PPO's clipped loss function. It uses the probability ratio between the current policy and the old model to cap how far one training step can move the model. A KL divergence penalty (a measure of how far the new model has drifted) keeps the policy close to a reference policy, usually the starting checkpoint, such as the supervised fine-tuning (SFT) model.

What GRPO does not remove is the scorer. DeepSeekMath scored its groups with a learned reward model initialised from DeepSeekMath-Base 7B. It sampled 64 outputs per question and trained on 144K chain-of-thought questions. The critic was gone; the reward model was still there.

Nathan Lambert's RLHF book notes that GRPO fits settings "when multiple completions to a given prompt is very natural." Maths and code are the obvious cases. Sampling eight or sixteen attempts at one problem is cheap, and checking them is often cheaper.

Sampling is the main difference in how the two algorithms spend their budget. In RLHF pipelines, PPO usually generates a single completion per prompt and relies on the critic for its baseline. GRPO needs a whole group for every prompt, and the memory freed by dropping the critic can go towards larger groups.

GRPO vs PPO at a glance

PPO GRPO
Baseline Learned value model (critic) Mean score of the sampled group
Models in memory during RL Policy, critic, reference, reward model Policy, reference, and a reward model if one is used
Samples per prompt Usually one A group: 16 in R1-Zero, 64 in DeepSeekMath, 8 by default in TRL
Where signal quality comes from Critic estimates plus reward Group composition plus reward

How Did DeepSeek-R1 Use GRPO Without a Reward Model?

The step that made GRPO famous came a year later, when DeepSeek used it to build reasoning capabilities through reinforcement learning alone. DeepSeek-R1 trained its R1-Zero model with GRPO and replaced the learned reward model with a rule-based reward function made of two checks. An accuracy reward checked the final answer against the ground truth. A format reward checked whether the reasoning sat inside the expected tags. The authors explain the choice directly: "the neural reward model may suffer from reward hacking in the large-scale reinforcement learning process." Reward hacking means the model finds unintended behaviors that score well without doing the task. R1-Zero sampled 16 outputs per question over 10,400 training steps.

This is the pattern now called reinforcement learning with verifiable rewards (RLVR). GRPO and RLVR are often mentioned together, but they are two different decisions. The key difference: GRPO removed the critic, and RLVR removed the learned reward model. Kili's earlier breakdown of DeepSeek R1 covers the full training pipeline.

R1 did not stay rule-based throughout. Its final RL stage brought back "reward models to capture human preferences in complex and nuanced scenarios," for helpfulness and harmlessness. The team also added a language-consistency reward, because pure RL produced reasoning that switched between languages. The paper was later peer reviewed and published in Nature. As Nature's news team reported, DeepSeek put the training cost of R1 at the equivalent of US$294,000.

Two larger deployments show where GRPO went next:

  • The Qwen3 technical report ran its reasoning RL stage with GRPO on 3,995 query–verifier pairs. On AIME'24 (a competition maths exam used as a benchmark), Qwen3-235B-A22B rose from 70.1 to 85.1 over 170 steps.
  • DeepSeek-V3.2 merged reasoning, agent and alignment training into a single GRPO stage, with a post-training computational cost above 10% of the pre-training budget. For general tasks it uses "a generative reward model where each prompt has its own rubrics for evaluation." Kili's DeepSeek V3.2 Data Story looks at what that rubric layer demands.

Across three generations of DeepSeek and Qwen models, the algorithm stayed recognisably the same. The scoring moved from a learned reward model, to rules, to per-prompt rubrics.

What Goes Wrong With Vanilla GRPO?

The 2024 equation has known problems, and most GRPO used in RL training today carries fixes.

Length bias. Liu et al. (2025) found that GRPO's per-response length normalisation "artificially increases response length (especially for incorrect outputs)." A wrong answer spread over more tokens gets a smaller penalty per token. So the model learns to ramble when it is lost. Their fix, Dr. GRPO, removes that normalisation along with the standard-deviation term.

Difficulty bias. Dividing by the group's standard deviation inflates advantages on prompts where nearly every answer agrees. Lambert's book describes the trade-off. Removing the term fixes the bias, but it weakens the signal from rare-success prompts. Those are the prompts where one correct answer in sixteen is exactly what you want to reinforce.

Instability at scale. ByteDance Seed's DAPO paper reports the result of its first attempt: "we achieved only 30 points on AIME" with naive GRPO, against DeepSeek's 47. DAPO made four changes and reached 50 points using half the training steps:

  • a looser upper clip, to prevent entropy collapse (the model becoming too certain to explore new answers)
  • dynamic sampling
  • a token-level loss
  • reward shaping for overlong answers

In practice, the "GRPO" in your LLM training library is likely not the 2024 GRPO. Take Hugging Face's TRL GRPOTrainer. At the time of writing, its default loss type is "dapo", its KL coefficient defaults to 0.0, and its default group size is 8. Check the configuration before comparing your results to a paper.

Two of DAPO's four fixes, dynamic sampling and overlong shaping, leave the optimiser alone. They change what goes into the group and how the group is scored.

Why Does GRPO Learn Nothing From Some Prompts?

Go back to the advantage formula. If all eight answers to a prompt are correct and each gets a reward of 1, the group mean is 1 and every advantage is 0. The same happens if all eight are wrong. The prompt consumed eight generations and a scoring pass, and moved the model nowhere.

DAPO states this plainly: when every output for a prompt "receive[s] the same reward 1, the resulting advantage for this group is zero." Its dynamic sampling step filters those prompts out. It keeps drawing until the batch is full of groups with mixed results.

Filtering at training time recovers compute. Choosing the training data well recovers more. Pikus, Tiwari and Ye (2025) tested this under a fixed budget: if you can only annotate 10% of a dataset, which 10% should it be? Training GRPO on the hardest examples produced gains of up to 47%, against 3–15% for easy ones. Only the hard-trained models improved on AIME 2025, a test set outside their training distribution. Their explanation is the formula again: "GRPO requires outcome variance to generate learning signals." Their recommendation reads like an annotation brief: "prioritize collecting and annotating examples where your base model struggles."

That study used small models and two task families, so treat the 47% as a direction, not a forecast. The direction is consistent with what larger labs do. The Qwen team chose its 3,995 reasoning-RL pairs because they were "learnable for the cold-start model" (the model after its first supervised stage) and "as challenging as possible."

For GRPO, the prompt set is a model-specific artefact. A prompt is useful only if the current policy sometimes gets it right and sometimes does not. That means:

  • Difficulty has to be measured against the model you are training. A difficulty label assigned once goes stale.
  • The useful band moves as the model improves, so yesterday's hard prompts become today's zero-advantage prompts.
  • Prompts need a reliable reference answer or rubric, or "sometimes right" cannot be measured at all.

What Happens When the Scorer Is Wrong?

In GRPO, the baseline is recomputed for every group from the scores themselves. Nothing independent checks them. Whatever the scorer gets wrong goes into the answer's reward and into the baseline that reward is compared with.

Scorers misjudge model outputs more often than "verifiable" suggests, and every error changes the reward signal.

Verifiers leak. A 2026 preregistered study looked at MBPP, a standard set of Python programming tasks. About 15% of tasks had natural false positives: weak tests that accept incorrect code. In the author's words, "the same wrong program passes the same weak tests on every rollout." During GRPO training, 8.82% of rewarded rollouts were false positives. The held-out capability loss was small (0.20 points), but training-side scores overstated progress. Cai et al. (2025) add that verifiers fail in both directions: they accept wrong answers and reject correct ones. Correcting for the false-negative rate alone improved RLVR.

Judges can be fooled. When a model is the scorer, the failure is easier to trigger. Zhao et al. (2025) tested leading LLM-as-a-judge systems, including GPT-o1 and Claude-4. Inputs as thin as a colon or the phrase "Thought process:" drew false-positive rewards.

A gain does not prove the reward was informative. Shao et al. (2025) trained Qwen2.5-Math-7B with GRPO on random rewards and still lifted its MATH-500 score by 21.4 points. Ground-truth rewards lifted it by 29.1. The effect largely disappeared on Llama3 and OLMo2. A rising curve on one model family tells you something about the model, and less than you would hope about your rewards.

The next article in this series, on RLVR, goes deeper into how verifiers get gamed. For GRPO, the lesson is shorter: audit the scorer before you trust the curve.

How Do You Score GRPO Groups When There Is No Verifier?

Maths answers can be checked. Clinical summaries, claim assessments and legal reasoning mostly cannot. For those tasks the choice is between a holistic judge score and a structured rubric.

The evidence favours structure. Rubrics as Rewards (Gunjal et al., 2025) turned a separate rubric for each prompt into GRPO rewards. Rubric rewards beat a judge giving one overall rating on a fixed scale (a Likert score). The relative improvement was up to 31% on HealthBench, a medical benchmark, and 7% on GPQA-Diamond, a graduate-level science benchmark. The authors work at Scale AI, a data vendor, so read the numbers as one team's result. The direction matches DeepSeek-V3.2's design, where every general-task prompt carries its own rubric.

A rubric is a set of human judgments written down in advance: what a correct answer must contain, what disqualifies it, how partial credit works. A GRPO run applies those judgments to thousands of groups, so it amplifies any ambiguity or error in the rubric. That puts three requirements on whoever builds the scoring layer:

  • Someone who can tell a correct answer from a plausible one in that domain has to write the rubric.
  • Reviewers have to apply it the same way. Two reviewers who read a criterion differently produce two different reward signals.
  • A sample of the automated scorer's verdicts needs regular comparison with an expert's, because judges drift and can be gamed.

This is the work Kili is built around. Kili is an expert-in-the-loop platform for AI evaluation and post-training. The domain expert judges directly, in configurable labeling and review workflows. Consensus measures how far experts agree with each other. Honeypot items check them against answers already known to be correct. Multi-step review adds further levels of review on top. For GRPO, that means producing and auditing the reference answers and expert judgments a scorer or rubric is calibrated against. It does not mean running the RL loop itself.

Does GRPO Work for Agents and Multi-Step Tasks?

Partly. GRPO assigns one advantage to a whole response. For a single maths answer, that is fine. For an agent taking twenty actions before a task succeeds or fails, it is coarse: every step gets the same credit.

GiGPO (Feng et al., NeurIPS 2025) names the problem: rewards that are "sparse or delayed" make "credit assignment across individual steps significantly more challenging." Its fix finds moments where different attempts reach the same environment state, and compares the actions taken from there. That gives step-level credit without an extra model. It beat GRPO by more than 12% on ALFWorld (simulated household tasks) and 9% on WebShop (simulated online shopping) at the same memory cost.

The data implication carries over. Step-level credit needs environments and tasks where intermediate states are identifiable and partial progress can be judged. Designing those checkpoints is, again, a scoring decision someone has to make.

Conclusion: GRPO Moves the Judgment, It Does Not Remove It

GRPO took a model out of reinforcement learning and put a statistic in its place. That was an engineering win. It means less memory and simpler code, and the algorithm has trained some of the strongest open reasoning models available.

The cost is that GRPO trusts its inputs. A group baseline is exactly as informative as the prompts that form the group and the scores assigned within it. Teams moving from maths to domains without clean verifiers should plan for the scoring layer to need more expert time than any other part of the stack. Budget for rubric design, reviewer agreement and scorer audits from the start, because in GRPO those are the training signal.

Resources

Original Papers and Model Reports

GRPO Variants and Fixes

Data Selection and Reward Quality

Tools and References

From Kili Technology

How Reliable Is the Signal in Your Post-Training Data?

Do your GRPO runs depend on rubrics, reference answers or expert review? Kili Technology helps domain experts produce and audit that scoring layer, with consensus and honeypot checks you configure per project.