Paper Methods
- Group Relative Policy Optimization (GRPO) without a Value Critic Network
- Intra-Group Z-Score Advantage Normalization across G Sampled Rollouts
- Unbiased Non-Negative Schulman Per-Token KL Divergence Penalty
- Outcome Supervision vs. Process Supervision Credit Assignment
Engineering Limitations
- •Requires generating G >= 8 rollouts per prompt, shifting the bottleneck from training memory to rollout generation throughput
- •Zero variance across all G group rollouts (e.g., all correct or all wrong) yields zero gradient signal for that prompt
- •Process Reward Models (PRMs) introduce step-boundary annotation overhead and potential reward hacking on intermediate reasoning steps
DeepSeekMath & GRPO Paper Breakdown
A mathematical and systems-level breakdown of DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (Shao et al., 2024, arXiv:2402.03300), the foundational paper that introduced Group Relative Policy Optimization (GRPO) and enabled the large-scale reinforcement learning pipeline behind DeepSeek-R1.
1. The PPO Critic Memory & Credit Assignment Bottleneck
Standard Reinforcement Learning from Human Feedback (RLHF) relies on Proximal Policy Optimization (PPO) (Schulman et al., 2017). To train an autoregressive policy pi_theta of parameter count P, PPO must simultaneously hold four large models across GPU High Bandwidth Memory (HBM):
Why the Value Critic Fails at Reasoning Scale
- Doubled Optimizer & Activation Memory: Because the Value Critic
V_psiis typically initialized as a full-size copy of the Policy Actor (e.g., 7B to 671B parameters) and trained concurrently via Generalized Advantage Estimation (GAE), it consumes ~50% of all trainable HBM (weights, gradients, and 8–12 bytes/param of Adam first/second moments). - Sparse Terminal Reward Variance: In mathematical reasoning and code synthesis, rewards are predominantly outcome-based—the verifier only knows whether the final boxed answer
\boxed{42}or unit test suite passed at the very last tokenTof a 16,384-token chain-of-thought (r_t = 0fort < T). Training a token-level criticV_psi(q, o_{<=t})to accurately predict terminal success across thousands of intermediate scratchpad tokens suffers from severe value estimation drift and high bias.
2. Group Relative Policy Optimization (GRPO) Formulation
GRPO completely eliminates the learned value function V_psi. Instead of estimating a baseline via a neural critic, GRPO samples a group of G distinct outputs {o_1, o_2, ..., o_G} from the old policy pi_{theta_old} for each question q, scores all G outputs, and uses the empirical group statistics as the advantage baseline.
The GRPO Surrogate Objective
For a query q ~ P(Q) and a group of G sampled completions {o_1, ..., o_G} ~ pi_{theta_old}(O | q), GRPO maximizes the clipped surrogate objective with an explicit per-token KL regularizer:
Where the per-token importance sampling ratio rho_{i,t} is:
Unbiased Non-Negative KL Divergence Estimator
Rather than folding the KL penalty into the reward signal r_t = R(q, o) - beta * log(...) (which complicates advantage normalization across varying sequence lengths), GRPO adds an explicit per-token regularizer directly to the loss using Schulman's unbiased, strictly non-negative estimator (k3 estimator):
Because x - log(x) - 1 >= 0 for all x > 0, this estimator is guaranteed to be non-negative and exhibits significantly lower variance than the naive Monte Carlo log-ratio -log(pi_ref / pi_theta), preventing negative KL spikes from destabilizing mixed-precision BF16/FP8 policy updates.
3. Outcome vs. Process Supervision in GRPO
DeepSeekMath formalizes two distinct credit assignment regimes within the GRPO framework:
3.1 Outcome Supervision GRPO
Each sampled completion o_i receives a single scalar reward r_i at the end of the sequence. The group vector r = [r_1, r_2, ..., r_G] is standardized using z-score normalization, and that normalized scalar is broadcast to every token t in trajectory i:
3.2 Process Supervision GRPO
When a Process Reward Model (PRM) scores intermediate reasoning steps o_i = {step_{i,1}, ..., step_{i,K_i}} at step-end token indices index(j), receives rewards R = {{r_1^{index(1)}, ..., r_1^{index(K_1)}}, ..., {r_G^{index(1)}, ..., r_G^{index(K_G)}}}, GRPO normalizes rewards across all steps in the group (r_hat_i^{index(j)}) and computes the token advantage as the cumulative future normalized step rewards:
Reference Implementation: Pure PyTorch GRPO Loss Kernel
4. Systems & Memory Comparison: PPO vs. GRPO
| Metric / Component | Standard PPO (7B Policy) | DeepSeekMath GRPO (7B Policy) | Systems Impact |
|---|---|---|---|
| Trainable Models in HBM | Actor (7B) + Critic (7B) | Actor (7B) only | ~50% reduction in trainable parameter & optimizer HBM |
| Adam Optimizer States (FP32) | ~112 GB (2 x 7B x 8B) | ~56 GB (1 x 7B x 8B) | Frees HBM for larger rollout batch sizes & longer CoT contexts |
| Advantage Baseline | Learned GAE (V_psi, lambda=0.95) | Group Empirical Mean/Std (G=64) | Zero value-network warm-up instability; exact Monte Carlo return |
| KL Penalty Injection | Folded into reward r_t | Direct per-token k3 loss term | Prevents KL reward distortion from skewing group advantage std(r) |
| GSM8K / MATH Accuracy (7B) | 82.9% / 46.8% (SFT/RFT baseline) | 88.2% / 51.7% | Surpasses open 70B models using purely group-relative RL |
5. Production Engineering Takeaways
- Inference-Bound RL Training: Because GRPO eliminates the critic backward pass and samples
G = 16to64trajectories per prompt, 70%–85% of wall-clock RL step time is spent in autoregressive rollout generation (vLLM/SGLangengines). Co-locating FP8 rollout engines with FSDP/Megatron training ranks via direct IPC weight synchronization is critical. - Dynamic Prompt Filtering (Zero-Variance Culling): When all
Grollouts for a prompt yield identical rewards (std(r) == 0, either 100% solved or 0% solved),A_{i,t} = 0for every token. Production GRPO pipelines dynamically filter out zero-variance groups before the backward pass to avoid wasted FLOPs. - Length Bias Mitigation (Dr. GRPO): Dividing by
|o_i|in the outer sum can subtly encourage longer incorrect responses (diluting negative advantages across more tokens) and shorter correct responses. Normalizing by a fixed maximum generation budgetT_maxeliminates length bias during long-CoT training.
