Research Paper Teardown
arXiv:2402.03300

DeepSeekMath & GRPO Paper Breakdown: Critic-Free Group Relative Policy Optimization for Reasoning LLMs

Definitive mathematical and distributed-systems teardown of DeepSeekMath (arXiv:2402.03300): eliminating the PPO value critic via Group Relative Policy Optimization (GRPO), intra-group advantage normalization, unbiased per-token KL divergence estimation, and outcome vs. process supervision.

18 min readVerified 2026-09-292 primary sourcesOriginal Paper
Technical paper breakdown illustration.

Paper Methods

  • Group Relative Policy Optimization (GRPO) without a Value Critic Network
  • Intra-Group Z-Score Advantage Normalization across G Sampled Rollouts
  • Unbiased Non-Negative Schulman Per-Token KL Divergence Penalty
  • Outcome Supervision vs. Process Supervision Credit Assignment

Engineering Limitations

  • •Requires generating G >= 8 rollouts per prompt, shifting the bottleneck from training memory to rollout generation throughput
  • •Zero variance across all G group rollouts (e.g., all correct or all wrong) yields zero gradient signal for that prompt
  • •Process Reward Models (PRMs) introduce step-boundary annotation overhead and potential reward hacking on intermediate reasoning steps

DeepSeekMath & GRPO Paper Breakdown

A mathematical and systems-level breakdown of DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (Shao et al., 2024, arXiv:2402.03300), the foundational paper that introduced Group Relative Policy Optimization (GRPO) and enabled the large-scale reinforcement learning pipeline behind DeepSeek-R1.


1. The PPO Critic Memory & Credit Assignment Bottleneck

Standard Reinforcement Learning from Human Feedback (RLHF) relies on Proximal Policy Optimization (PPO) (Schulman et al., 2017). To train an autoregressive policy pi_theta of parameter count P, PPO must simultaneously hold four large models across GPU High Bandwidth Memory (HBM):

text(9 lines)
1+-------------------------------------------------------------------------+
2| Standard PPO Distributed HBM Footprint (4 Full-Scale Models) |
3+-------------------------------------------------------------------------+
4| 1. Policy Actor (pi_theta) : Trainable weights + FP32 Adam states |
5| 2. Value Critic (V_psi) : Trainable weights + FP32 Adam states |
6| 3. Reference Policy (pi_ref) : Frozen FP16/BF16 weights |
7| 4. Reward Model (R_phi) : Frozen FP16/BF16 weights (or Verifier)|
8+-------------------------------------------------------------------------+

Why the Value Critic Fails at Reasoning Scale

  1. Doubled Optimizer & Activation Memory: Because the Value Critic V_psi is typically initialized as a full-size copy of the Policy Actor (e.g., 7B to 671B parameters) and trained concurrently via Generalized Advantage Estimation (GAE), it consumes ~50% of all trainable HBM (weights, gradients, and 8–12 bytes/param of Adam first/second moments).
  2. Sparse Terminal Reward Variance: In mathematical reasoning and code synthesis, rewards are predominantly outcome-based—the verifier only knows whether the final boxed answer \boxed{42} or unit test suite passed at the very last token T of a 16,384-token chain-of-thought (r_t = 0 for t < T). Training a token-level critic V_psi(q, o_{<=t}) to accurately predict terminal success across thousands of intermediate scratchpad tokens suffers from severe value estimation drift and high bias.

2. Group Relative Policy Optimization (GRPO) Formulation

GRPO completely eliminates the learned value function V_psi. Instead of estimating a baseline via a neural critic, GRPO samples a group of G distinct outputs {o_1, o_2, ..., o_G} from the old policy pi_{theta_old} for each question q, scores all G outputs, and uses the empirical group statistics as the advantage baseline.

text(6 lines)
1 +---> o_1 ---> r_1 ---+
2 | |
3Prompt q ---> pi_old ----+---> o_2 ---> r_2 ---+---> Group Normalization
4 | ... | A_i = (r_i - mean(r)) / std(r)
5 +---> o_G ---> r_G ---+

The GRPO Surrogate Objective

For a query q ~ P(Q) and a group of G sampled completions {o_1, ..., o_G} ~ pi_{theta_old}(O | q), GRPO maximizes the clipped surrogate objective with an explicit per-token KL regularizer:

text(7 lines)
1J_GRPO(theta) = E_{q, {o_i}} [
2 (1 / G) * sum_{i=1..G} (1 / |o_i|) * sum_{t=1..|o_i|} (
3 min( rho_{i,t} * A_{i,t}, clip(rho_{i,t}, 1 - eps, 1 + eps) * A_{i,t} )
4 - beta * D_KL( pi_theta || pi_ref )
5 )
6]

Where the per-token importance sampling ratio rho_{i,t} is:

text(2 lines)
1rho_{i,t} = pi_theta(o_{i,t} | q, o_{i,<t}) / pi_{theta_old}(o_{i,t} | q, o_{i,<t})

Unbiased Non-Negative KL Divergence Estimator

Rather than folding the KL penalty into the reward signal r_t = R(q, o) - beta * log(...) (which complicates advantage normalization across varying sequence lengths), GRPO adds an explicit per-token regularizer directly to the loss using Schulman's unbiased, strictly non-negative estimator (k3 estimator):

text(4 lines)
1D_KL( pi_theta || pi_ref ) = ( pi_ref(o_{i,t} | q, o_{i,<t}) / pi_theta(o_{i,t} | q, o_{i,<t}) )
2 - log( pi_ref(o_{i,t} | q, o_{i,<t}) / pi_theta(o_{i,t} | q, o_{i,<t}) )
3 - 1

Because x - log(x) - 1 >= 0 for all x > 0, this estimator is guaranteed to be non-negative and exhibits significantly lower variance than the naive Monte Carlo log-ratio -log(pi_ref / pi_theta), preventing negative KL spikes from destabilizing mixed-precision BF16/FP8 policy updates.


3. Outcome vs. Process Supervision in GRPO

DeepSeekMath formalizes two distinct credit assignment regimes within the GRPO framework:

3.1 Outcome Supervision GRPO

Each sampled completion o_i receives a single scalar reward r_i at the end of the sequence. The group vector r = [r_1, r_2, ..., r_G] is standardized using z-score normalization, and that normalized scalar is broadcast to every token t in trajectory i:

text(2 lines)
1A_{i,t} = A_i = ( r_i - mean(r_1, ..., r_G) ) / ( std(r_1, ..., r_G) + eps )

3.2 Process Supervision GRPO

When a Process Reward Model (PRM) scores intermediate reasoning steps o_i = {step_{i,1}, ..., step_{i,K_i}} at step-end token indices index(j), receives rewards R = {{r_1^{index(1)}, ..., r_1^{index(K_1)}}, ..., {r_G^{index(1)}, ..., r_G^{index(K_G)}}}, GRPO normalizes rewards across all steps in the group (r_hat_i^{index(j)}) and computes the token advantage as the cumulative future normalized step rewards:

text(2 lines)
1A_{i,t} = sum_{index(j) >= t} r_hat_i^{index(j)}

Reference Implementation: Pure PyTorch GRPO Loss Kernel

python(39 lines)
1import torch
2
3def compute_grpo_loss(
4 logprobs_new: torch.Tensor, # [G, T] log pi_theta(o_{i,t} | q, o_{i,<t})
5 logprobs_old: torch.Tensor, # [G, T] log pi_{theta_old}(o_{i,t} | q, o_{i,<t})
6 logprobs_ref: torch.Tensor, # [G, T] log pi_ref(o_{i,t} | q, o_{i,<t})
7 rewards: torch.Tensor, # [G] terminal outcome rewards per rollout
8 mask: torch.Tensor, # [G, T] binary completion mask (1 for valid tokens)
9 clip_eps: float = 0.2,
10 beta_kl: float = 0.04,
11) -> dict[str, torch.Tensor]:
12 # 1. Intra-group z-score advantage normalization: [G, 1]
13 mean_r = rewards.mean()
14 std_r = rewards.std(unbiased=False)
15 advantages = ((rewards - mean_r) / (std_r + -8)).unsqueeze(1) # Broadcast across T
16
17 # 2. Per-token importance sampling ratio rho_{i,t}
18 ratio = torch.exp(logprobs_new - logprobs_old)
19
20 # 3. Clipped PPO-style surrogate objective
21 surr1 = ratio * advantages
22 surr2 = torch.clamp(ratio, 1.0 - clip_eps, 1.0 + clip_eps) * advantages
23 policy_obj = torch.minimum(surr1, surr2)
24
25 # 4. Unbiased non-negative per-token KL penalty (Schulman k3 estimator)
26 log_ratio_ref = logprobs_ref - logprobs_new
27 kl_div = torch.exp(log_ratio_ref) - log_ratio_ref - 1.0
28
29 # 5. Sequence-length normalized group average loss
30 per_token_loss = -(policy_obj - beta_kl * kl_div) * mask
31 seq_lengths = mask.sum(dim=1).clamp(min=1.0)
32 loss = (per_token_loss.sum(dim=1) / seq_lengths).mean()
33
34 return {
35 "loss": loss,
36 "mean_kl": (kl_div * mask).sum() / mask.sum().clamp(min=1.0),
37 "advantages": advantages.squeeze(1),
38 }
19 lines hidden

4. Systems & Memory Comparison: PPO vs. GRPO

| Metric / Component | Standard PPO (7B Policy) | DeepSeekMath GRPO (7B Policy) | Systems Impact | |---|---|---|---| | Trainable Models in HBM | Actor (7B) + Critic (7B) | Actor (7B) only | ~50% reduction in trainable parameter & optimizer HBM | | Adam Optimizer States (FP32) | ~112 GB (2 x 7B x 8B) | ~56 GB (1 x 7B x 8B) | Frees HBM for larger rollout batch sizes & longer CoT contexts | | Advantage Baseline | Learned GAE (V_psi, lambda=0.95) | Group Empirical Mean/Std (G=64) | Zero value-network warm-up instability; exact Monte Carlo return | | KL Penalty Injection | Folded into reward r_t | Direct per-token k3 loss term | Prevents KL reward distortion from skewing group advantage std(r) | | GSM8K / MATH Accuracy (7B) | 82.9% / 46.8% (SFT/RFT baseline) | 88.2% / 51.7% | Surpasses open 70B models using purely group-relative RL |


5. Production Engineering Takeaways

  1. Inference-Bound RL Training: Because GRPO eliminates the critic backward pass and samples G = 16 to 64 trajectories per prompt, 70%–85% of wall-clock RL step time is spent in autoregressive rollout generation (vLLM / SGLang engines). Co-locating FP8 rollout engines with FSDP/Megatron training ranks via direct IPC weight synchronization is critical.
  2. Dynamic Prompt Filtering (Zero-Variance Culling): When all G rollouts for a prompt yield identical rewards (std(r) == 0, either 100% solved or 0% solved), A_{i,t} = 0 for every token. Production GRPO pipelines dynamically filter out zero-variance groups before the backward pass to avoid wasted FLOPs.
  3. Length Bias Mitigation (Dr. GRPO): Dividing by |o_i| in the outer sum can subtly encourage longer incorrect responses (diluting negative advantages across more tokens) and shorter correct responses. Normalizing by a fixed maximum generation budget T_max eliminates length bias during long-CoT training.