Architecture
Post-training has evolved from supervised instruction tuning and preference modeling into reinforcement learning with verifiable reward signals (RLVR). In frontier reasoning architectures such as DeepSeekMath and DeepSeek-R1, reinforcement learning acts as an exploratory search engine over long chain-of-thought (CoT) trajectories, incentivizing self-correction, backtracking, and algorithmic deduction.
Traditional Reinforcement Learning from Human Feedback (RLHF) utilizes Proximal Policy Optimization (PPO) structured around an Actor-Critic architecture. While effective for conversational alignment, standard PPO introduces catastrophic bottlenecks when applied to long-context reasoning models:
- VRAM Memory Bloat: PPO requires maintaining four distinct large language models concurrently: the trainable Actor (
\pi_\theta), the trainable Critic (V_\phi), the frozen Reference Model (\pi_{\text{ref}}), and the frozen Reward Model (R_\psi). For 67B or 671B Mixture-of-Experts architectures, hosting the Critic's parameters, optimizer states, and activations consumes up to 50% of the entire GPU cluster's High-Bandwidth Memory (HBM). - Critic Instability & Value Drift: A value network must predict expected future scalar rewards across token generation steps. Because reasoning problems (Olympiad mathematics, code synthesis) exhibit sparse, binary outcomes (correct answer or wrong answer at token 8,192), training a stable token-level value function
V_\phi(s_t)is notoriously fragile and prone to value estimation collapse. - Reward Hacking in Neural Verifiers: Learned neural reward models are easily exploited by policies generating verbose, pseudo-authoritative tokens that score high on subjective alignment heuristics without solving the underlying logical problem.
Group Relative Policy Optimization (GRPO) resolves these constraints by completely discarding the Critic network. Instead of evaluating state values against a neural baseline V_\phi(s), GRPO samples a group of G candidate reasoning paths \{o_1, o_2, \dots, o_G\} for each input prompt q, scores each path using deterministic, rule-based reward functions, and normalizes advantages strictly relative to the group mean and standard deviation.
Mathematical Foundations: Actor-Critic PPO vs. Critic-Free GRPO
To understand the algorithmic divergence between PPO and GRPO, we trace their respective objective functions, advantage estimation formulations, and divergence constraints.
1. PPO Clipped Surrogate Objective with Generalized Advantage Estimation (GAE)
In standard PPO, policy parameters \theta are updated by maximizing the clipped surrogate objective over trajectories generated by an older policy checkpoint \pi_{\theta_{\text{old}}}:
where the probability ratio r_t(\theta) is defined as:
The advantage \hat{A}_t is computed using Generalized Advantage Estimation (\text{GAE}(\gamma, \lambda)) relying on the scalar value predictions of the Critic network V_\phi:
The Critic parameters \phi are simultaneously trained by minimizing mean squared error against empirical discounted returns:
In multi-billion parameter reasoning tasks, maintaining and synchronizing the gradient updates for V_\phi alongside \pi_\theta doubles the communication overhead during distributed backward passes.
2. GRPO Formulation: Group Relative Advantage Normalization
GRPO obviates the need for V_\phi by computing empirical baselines directly across sampled completions. For each query q \sim P(Q), the policy generates a group of G distinct outputs:
Each completion o_i receives a trajectory-level scalar reward r_i \in \mathbb{R} computed by deterministic verifiers. The empirical group mean \mu_r and standard deviation \sigma_r are calculated as:
The normalized advantage for completion o_i is computed without any learned value baseline:
where \epsilon > 0 is a small numerical stabilization constant. Notice that \hat{A}_i is a sequence-level scalar applied uniformly across all tokens t \in \{1, \dots, |o_i|\} within completion o_i.
The complete GRPO policy objective is defined as:
Unbiased Non-Negative Schulman k_3 KL Divergence Estimator
To prevent policy drift away from the foundational capabilities of the reference model \pi_{\text{ref}}, reinforcement learning objectives penalize Kullback-Leibler (KL) divergence.
In standard implementations, the naïve k_1 estimator evaluates the log-ratio:
Under stochastic sample approximations, D_{\text{KL}}^{k_1} can evaluate to negative values when \pi_\theta(x) < \pi_{\text{ref}}(x), introducing high-variance gradient destabilization during long reasoning generations.
GRPO employs the Schulman k_3 unbiased, non-negative KL estimator:
Let the likelihood ratio be denoted by r = \frac{\pi_{\text{ref}}(t)}{\pi_\theta(t)} \in (0, \infty). We analyze the scalar function:
- First derivative:
f'(r) = 1 - \frac{1}{r}. Settingf'(r) = 0yields a single stationary point atr = 1. - Second derivative:
f''(r) = \frac{1}{r^2} > 0for allr \in (0, \infty), confirming thatf(r)is strictly convex. - Global minimum: At
r = 1,f(1) = 1 - 0 - 1 = 0.
Therefore, for all token probability distributions where r > 0:
This property guarantees that the policy is never inadvertently rewarded for drifting away from the reference model, stabilizing multi-turn reasoning exploration.
GPU Memory Budgeting & Actor-Critic Allocations
In large-scale post-training, GPU High-Bandwidth Memory (HBM) is divided across model weights, gradient buffers, optimizer states (typically AdamW), and activation memory.
1. Per-Model VRAM Requirements in 16-bit Precision
Under ZeRO-3 / FSDP distributed sharding, each model component requires:
- Trainable Parameters (Actor & Critic):
- Model Weights (BF16): 2 bytes/parameter
- Gradient Buffers (BF16): 2 bytes/parameter
- AdamW Optimizer States (FP32 master weights, 1st momentum, 2nd variance): 12 bytes/parameter
- Total Trainable Model State: 16 bytes/parameter
- Frozen Parameters (Reference & Neural Reward Model):
- Model Weights (BF16 or FP8): 2 bytes/parameter (or 1 byte in FP8)
- Gradient & Optimizer: 0 bytes
- Total Frozen Model State: 2 bytes/parameter
2. VRAM Comparison: 4-Model PPO vs. 2-Model GRPO
| Architecture Scale | Active Parameters | 4-Model PPO Footprint (Actor + Critic + Ref + Reward) | 2-Model GRPO Footprint (Actor + Ref + Verifier Sandbox) | Net VRAM Reduction | |---|---|---|---|---| | Small Dense (7B) | 7.24 Billion | 260.6 GB VRAM | 130.3 GB VRAM | 50.0% Savings | | Mid-Tier (67B) | 67.1 Billion | 2,415.6 GB VRAM (31x H100 80GB) | 1,207.8 GB VRAM (16x H100 80GB) | 50.0% Savings | | Frontier MoE (671B Total / 37B Active) | 671.0 Billion | 24,156.0 GB VRAM (302x H100 80GB) | 12,078.0 GB VRAM (151x H100 80GB) | 50.0% Savings |
By eliminating the 16 bytes/parameter required for the Critic and the 2 bytes/parameter required for a neural Reward Model, GRPO saves precisely 50% of model-state VRAM across the training cluster. This surplus memory is redirected to extended sequence KV-caches and activation checkpointing buffers, allowing reasoning models to scale generation horizons from 4,096 tokens to 32,768 tokens without triggering out-of-memory (OOM) faults.
Verifiable Rule-Based Reward Engineering & The Zero-Variance Clamp
Neural reward models are susceptible to Goodhart's Law: when a proxy reward metric is optimized aggressively, it ceases to be a valid measure of competence. In reasoning domains, RLVR enforces strict, deterministic verifiers.
The Zero-Variance Optimization Collapse (Failure Mode)
In group-relative advantage normalization:
Consider a batch where all G completions for a difficult mathematical problem fail completely:
Similarly, if all G completions solve an easy problem identically:
When \sigma_r = 0, all candidate completions performed identically. There is zero empirical contrastive signal to indicate which tokens contributed to success or failure.
If \epsilon is very small (e.g., 10^{-8}), floating-point rounding errors or subnormal numbers cause \hat{A}_i to produce massive, erratic gradients. Conversely, if \epsilon is moderate, each completion receives \hat{A}_i = \frac{0}{\epsilon} = 0, which is mathematically correct but prone to floating-point drift.
More dangerously, if completions achieve identical partial format rewards (r_i = 0.1) without correct answers (r_{\text{acc}} = 0), calculating unconstrained advantages can result in updating policy weights toward repetitive thinking patterns.
The Zero-Variance Clamp Guardrail
To maintain strict mathematical stability across distributed ranks, the advantage calculation must enforce an explicit zero-variance guard:
When \sigma_r < 10^{-8}, setting \hat{A}_i = 0.0 forces the clipped surrogate policy gradient \nabla_\theta \mathcal{J}(\theta) to zero for that group prompt, effectively bypassing parameter updates on non-informative rollouts.
Decisions
| Decision | Required evidence | Review trigger |
|---|---|---|
| Select GRPO over PPO for reasoning post-training | Verification that target domain admits deterministic rule-based verifiers (math, code, structured JSON). | Migrating post-training pipelines from conversational chat to mathematical/algorithmic reasoning. |
| Set Group Size to G = 8 for dense models and G = 16 for MoE | Empirical validation that group variance \sigma_r^2 > 0 on at least 70% of training prompts. | Observational telemetry indicating > 40\% zero-variance clamp engagement. |
| Enforce Schulman k_3 KL penalty with coefficient \beta \in [0.01, 0.05] | KL divergence monitoring showing strictly non-negative step metrics and absence of policy collapse. | Mean sequence length exploding by > 2.5\times baseline without benchmark accuracy gains. |
| Isolate Code/Math Verifiers in gVisor / Firecracker microVMs | Benchmark execution latency under 250ms per test suite with complete network namespace isolation. | Deploying unvetted multi-turn code synthesis RL pipelines to distributed training nodes. |
Alternatives and trade-offs
- Direct Preference Optimization (DPO): While DPO is computationally lightweight (requiring no online sampling during parameter updates), it operates over static, pre-collected preference pairs. DPO cannot discover novel self-correction strategies or multi-step reasoning trajectories that were absent from the offline dataset.
- Actor-Critic PPO: Remains mandatory for subjective tasks (creative writing, general conversational tone, nuanced persona alignment) where deterministic verification is impossible and a learned reward model is required.
- Group Relative Policy Optimization (GRPO): Optimal for objective tasks with deterministic ground-truth verification. Eliminating the Critic cuts cluster VRAM requirements in half while group normalization provides robust contrastive signals for policy improvement.
Failure modes
- The Zero-Variance Optimization Plateau: Occurs when training prompts are either too simple (100% group success) or impossibly complex (0% group success). With
\sigma_r = 0, advantages evaluate to zero, causing gradient starvation and halting learning progression. - Runaway Chain-of-Thought Reflection Loops: If length penalties are insufficiently calibrated, the model discovers that generating circular reflective markers (e.g., "Wait, let me rethink this...") delays completion while preserving format rewards, resulting in maximum token limit saturation.
- Language Divergence & Code Switching: During pure RLVR on English reasoning datasets, models initialized from multilingual checkpoints may spontaneously switch languages mid-trajectory if token probability distributions in another language happen to yield lower entropy. Mitigated by adding language consistency verification to
r_{\text{format}}. - Format Reward Exploitation: The policy generates thousands of empty XML tags or injects multiple closing tags to fool naive regex parsers into assigning maximum format points without executing valid reasoning steps.
Operational checklist
- [ ] All reward evaluation pipelines execute inside hardened, network-isolated sandboxes with execution timeouts under 500ms.
- [ ] Group advantage normalization includes an explicit zero-variance guard (
\sigma_r < 10^{-8} \implies \hat{A}_i = 0.0). - [ ] Schulman
k_3KL divergence formulation is validated to remain non-negative across all distributed training ranks. - [ ] Group rollout generation utilizes PagedAttention to eliminate KV-cache memory fragmentation during variable-length rollouts.
- [ ] Training telemetry monitors the Zero-Variance Engagement Rate (
\% \text{ of batches clamped}); if rate exceeds 35%, dynamic curriculum rebalancing is triggered. - [ ] Sequence length limits include safety margins to prevent distributed all-gather deadlocks when multiple completions hit max tokens.
Connected practice
- Labs: /labs/grpo-reasoning
- System breakdowns: /systems/inside-deepseek-v3-r1, /systems/inside-vllm
- Paper breakdowns: /papers/deepseekmath-grpo-paper
Sources
deepseekmath-grpo-2024deepseek-r1-paper
