Reference architecture

RLHF, PPO & GRPO: Reasoning Alignment & Policy Optimization Post-Training

Mathematical foundations of Actor-Critic PPO vs. critic-free Group Relative Policy Optimization (GRPO), verifiable rule-based reward engineering, and actor-critic memory budgeting for reasoning models.

22 minVerified 2026-09-292 primary sources
A governed production AI reference architecture with observable, secured service boundaries.

Architecture

Post-training has evolved from supervised instruction tuning and preference modeling into reinforcement learning with verifiable reward signals (RLVR). In frontier reasoning architectures such as DeepSeekMath and DeepSeek-R1, reinforcement learning acts as an exploratory search engine over long chain-of-thought (CoT) trajectories, incentivizing self-correction, backtracking, and algorithmic deduction.

Prompt sampling from reasoning task dataset
Group rollout generation (G completions per prompt)
Deterministic rule-based and format reward scoring
Group mean and variance normalization
Schulman k3 unbiased KL penalty estimation
Clipped surrogate policy gradient backpropagation
Conceptual teaching model synthesized from:DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeek-R1 Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Traditional Reinforcement Learning from Human Feedback (RLHF) utilizes Proximal Policy Optimization (PPO) structured around an Actor-Critic architecture. While effective for conversational alignment, standard PPO introduces catastrophic bottlenecks when applied to long-context reasoning models:

  1. VRAM Memory Bloat: PPO requires maintaining four distinct large language models concurrently: the trainable Actor (\pi_\theta), the trainable Critic (V_\phi), the frozen Reference Model (\pi_{\text{ref}}), and the frozen Reward Model (R_\psi). For 67B or 671B Mixture-of-Experts architectures, hosting the Critic's parameters, optimizer states, and activations consumes up to 50% of the entire GPU cluster's High-Bandwidth Memory (HBM).
  2. Critic Instability & Value Drift: A value network must predict expected future scalar rewards across token generation steps. Because reasoning problems (Olympiad mathematics, code synthesis) exhibit sparse, binary outcomes (correct answer or wrong answer at token 8,192), training a stable token-level value function V_\phi(s_t) is notoriously fragile and prone to value estimation collapse.
  3. Reward Hacking in Neural Verifiers: Learned neural reward models are easily exploited by policies generating verbose, pseudo-authoritative tokens that score high on subjective alignment heuristics without solving the underlying logical problem.

Group Relative Policy Optimization (GRPO) resolves these constraints by completely discarding the Critic network. Instead of evaluating state values against a neural baseline V_\phi(s), GRPO samples a group of G candidate reasoning paths \{o_1, o_2, \dots, o_G\} for each input prompt q, scores each path using deterministic, rule-based reward functions, and normalizes advantages strictly relative to the group mean and standard deviation.

Model weight sharding across cluster ranks
Critic model activation and optimizer memory footprint
Reference policy KV cache and frozen weight caching
Reward model inference pipeline allocations
Critic-free consolidation in GRPO
Available VRAM redirected to extended reasoning context
Conceptual teaching model synthesized from:DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language ModelsDeepSeek-R1 Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Mathematical Foundations: Actor-Critic PPO vs. Critic-Free GRPO

To understand the algorithmic divergence between PPO and GRPO, we trace their respective objective functions, advantage estimation formulations, and divergence constraints.

1. PPO Clipped Surrogate Objective with Generalized Advantage Estimation (GAE)

In standard PPO, policy parameters \theta are updated by maximizing the clipped surrogate objective over trajectories generated by an older policy checkpoint \pi_{\theta_{\text{old}}}:

text(2 lines)
1\mathcal{L}_{\text{PPO}}(\theta) = \hat{\mathbb{E}}_t \left[ \min\left( r_t(\theta) \hat{A}_t, \, \text{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon) \hat{A}_t \right) \right]

where the probability ratio r_t(\theta) is defined as:

text(2 lines)
1r_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\theta_{\text{old}}}(a_t \mid s_t)}

The advantage \hat{A}_t is computed using Generalized Advantage Estimation (\text{GAE}(\gamma, \lambda)) relying on the scalar value predictions of the Critic network V_\phi:

text(2 lines)
1\hat{A}_t^{\text{GAE}} = \sum_{l=0}^{\infty} (\gamma \lambda)^l \delta_{t+l}^V, \quad \text{where} \quad \delta_t^V = R_t + \gamma V_\phi(s_{t+1}) - V_\phi(s_t)

The Critic parameters \phi are simultaneously trained by minimizing mean squared error against empirical discounted returns:

text(2 lines)
1\mathcal{L}_{\text{Critic}}(\phi) = \hat{\mathbb{E}}_t \left[ \left( V_\phi(s_t) - R_t^{\text{target}} \right)^2 \right]

In multi-billion parameter reasoning tasks, maintaining and synchronizing the gradient updates for V_\phi alongside \pi_\theta doubles the communication overhead during distributed backward passes.

2. GRPO Formulation: Group Relative Advantage Normalization

GRPO obviates the need for V_\phi by computing empirical baselines directly across sampled completions. For each query q \sim P(Q), the policy generates a group of G distinct outputs:

text(2 lines)
1\mathcal{O}_q = \{o_1, o_2, \dots, o_G\}, \quad o_i \sim \pi_{\theta_{\text{old}}}(\cdot \mid q)

Each completion o_i receives a trajectory-level scalar reward r_i \in \mathbb{R} computed by deterministic verifiers. The empirical group mean \mu_r and standard deviation \sigma_r are calculated as:

text(2 lines)
1\mu_r = \frac{1}{G} \sum_{i=1}^G r_i, \qquad \sigma_r = \sqrt{\frac{1}{G} \sum_{i=1}^G (r_i - \mu_r)^2}

The normalized advantage for completion o_i is computed without any learned value baseline:

text(2 lines)
1\hat{A}_i = \frac{r_i - \mu_r}{\sigma_r + \epsilon}

where \epsilon > 0 is a small numerical stabilization constant. Notice that \hat{A}_i is a sequence-level scalar applied uniformly across all tokens t \in \{1, \dots, |o_i|\} within completion o_i.

The complete GRPO policy objective is defined as:

text(2 lines)
1\mathcal{J}_{\text{GRPO}}(\theta) = \mathbb{E}_{\substack{q \sim P(Q) \\ \{o_i\}_{i=1}^G \sim \pi_{\theta_{\text{old}}}}} \left[ \frac{1}{G} \sum_{i=1}^G \frac{1}{|o_i|} \sum_{t=1}^{|o_i|} \left\{ \min\left( \frac{\pi_\theta(o_{i,t} \mid q, o_{i,<t})}{\pi_{\theta_{\text{old}}}(o_{i,t} \mid q, o_{i,<t})} \hat{A}_i, \, \text{clip}\left(\frac{\pi_\theta(o_{i,t} \mid q, o_{i,<t})}{\pi_{\theta_{\text{old}}}(o_{i,t} \mid q, o_{i,<t})}, 1-\epsilon, 1+\epsilon\right) \hat{A}_i \right) - \beta D_{\text{KL}}(\pi_\theta \parallel \pi_{\text{ref}}) \right\} \right]

Unbiased Non-Negative Schulman k_3 KL Divergence Estimator

To prevent policy drift away from the foundational capabilities of the reference model \pi_{\text{ref}}, reinforcement learning objectives penalize Kullback-Leibler (KL) divergence.

In standard implementations, the naïve k_1 estimator evaluates the log-ratio:

text(2 lines)
1D_{\text{KL}}^{k_1}(t) = \log\left(\frac{\pi_{\text{ref}}(o_{i,t} \mid q, o_{i,<t})}{\pi_\theta(o_{i,t} \mid q, o_{i,<t})}\right) = -\log r_t

Under stochastic sample approximations, D_{\text{KL}}^{k_1} can evaluate to negative values when \pi_\theta(x) < \pi_{\text{ref}}(x), introducing high-variance gradient destabilization during long reasoning generations.

GRPO employs the Schulman k_3 unbiased, non-negative KL estimator:

text(2 lines)
1D_{\text{KL}}^{k_3}(t) = \frac{\pi_{\text{ref}}(o_{i,t} \mid q, o_{i,<t})}{\pi_\theta(o_{i,t} \mid q, o_{i,<t})} - \log\left(\frac{\pi_{\text{ref}}(o_{i,t} \mid q, o_{i,<t})}{\pi_\theta(o_{i,t} \mid q, o_{i,<t})}\right) - 1

Let the likelihood ratio be denoted by r = \frac{\pi_{\text{ref}}(t)}{\pi_\theta(t)} \in (0, \infty). We analyze the scalar function:

text(2 lines)
1f(r) = r - \log(r) - 1
  1. First derivative: f'(r) = 1 - \frac{1}{r}. Setting f'(r) = 0 yields a single stationary point at r = 1.
  2. Second derivative: f''(r) = \frac{1}{r^2} > 0 for all r \in (0, \infty), confirming that f(r) is strictly convex.
  3. Global minimum: At r = 1, f(1) = 1 - 0 - 1 = 0.

Therefore, for all token probability distributions where r > 0:

text(2 lines)
1D_{\text{KL}}^{k_3}(t) \ge 0 \quad \text{strictly point-wise}

This property guarantees that the policy is never inadvertently rewarded for drifting away from the reference model, stabilizing multi-turn reasoning exploration.


GPU Memory Budgeting & Actor-Critic Allocations

In large-scale post-training, GPU High-Bandwidth Memory (HBM) is divided across model weights, gradient buffers, optimizer states (typically AdamW), and activation memory.

1. Per-Model VRAM Requirements in 16-bit Precision

Under ZeRO-3 / FSDP distributed sharding, each model component requires:

  • Trainable Parameters (Actor & Critic):
    • Model Weights (BF16): 2 bytes/parameter
    • Gradient Buffers (BF16): 2 bytes/parameter
    • AdamW Optimizer States (FP32 master weights, 1st momentum, 2nd variance): 12 bytes/parameter
    • Total Trainable Model State: 16 bytes/parameter
  • Frozen Parameters (Reference & Neural Reward Model):
    • Model Weights (BF16 or FP8): 2 bytes/parameter (or 1 byte in FP8)
    • Gradient & Optimizer: 0 bytes
    • Total Frozen Model State: 2 bytes/parameter

2. VRAM Comparison: 4-Model PPO vs. 2-Model GRPO

| Architecture Scale | Active Parameters | 4-Model PPO Footprint (Actor + Critic + Ref + Reward) | 2-Model GRPO Footprint (Actor + Ref + Verifier Sandbox) | Net VRAM Reduction | |---|---|---|---|---| | Small Dense (7B) | 7.24 Billion | 260.6 GB VRAM | 130.3 GB VRAM | 50.0% Savings | | Mid-Tier (67B) | 67.1 Billion | 2,415.6 GB VRAM (31x H100 80GB) | 1,207.8 GB VRAM (16x H100 80GB) | 50.0% Savings | | Frontier MoE (671B Total / 37B Active) | 671.0 Billion | 24,156.0 GB VRAM (302x H100 80GB) | 12,078.0 GB VRAM (151x H100 80GB) | 50.0% Savings |

By eliminating the 16 bytes/parameter required for the Critic and the 2 bytes/parameter required for a neural Reward Model, GRPO saves precisely 50% of model-state VRAM across the training cluster. This surplus memory is redirected to extended sequence KV-caches and activation checkpointing buffers, allowing reasoning models to scale generation horizons from 4,096 tokens to 32,768 tokens without triggering out-of-memory (OOM) faults.


Verifiable Rule-Based Reward Engineering & The Zero-Variance Clamp

Neural reward models are susceptible to Goodhart's Law: when a proxy reward metric is optimized aggressively, it ceases to be a valid measure of competence. In reasoning domains, RLVR enforces strict, deterministic verifiers.

text(23 lines)
1+-------------------------------------------------------------------------------+
2| COMPLETION EVALUATION PIPELINE |
3+-------------------------------------------------------------------------------+
4 Input Completion (Rollout o_i)
5 |
6 v
7 [Format Verification Engine]
8 |---> Validates <think>...</think><answer>...</answer> XML tags
9 |---> Computes r_format in [0.0, 0.3]
10 |
11 v
12 [Domain Verifier Sandbox]
13 |---> Symbolic Math: SymPy algebraic equivalence against ground truth
14 |---> Code Execution: Isolated Docker/gVisor test suite execution
15 |---> Computes r_accuracy in {0.0, 1.0}
16 |
17 v
18 [Length Penalty Engine]
19 |---> Identifies uninformative loop recursion (r_len in [-0.2, 0.0])
20 |
21 v
22 Total Scalar Reward: r_i = r_format + r_accuracy + r_len

The Zero-Variance Optimization Collapse (Failure Mode)

In group-relative advantage normalization:

text(2 lines)
1\hat{A}_i = \frac{r_i - \mu_r}{\sigma_r + \epsilon}

Consider a batch where all G completions for a difficult mathematical problem fail completely:

text(2 lines)
1r_1 = 0, \, r_2 = 0, \, \dots, \, r_G = 0 \implies \mu_r = 0, \quad \sigma_r = 0

Similarly, if all G completions solve an easy problem identically:

text(2 lines)
1r_1 = 1, \, r_2 = 1, \, \dots, \, r_G = 1 \implies \mu_r = 1, \quad \sigma_r = 0

When \sigma_r = 0, all candidate completions performed identically. There is zero empirical contrastive signal to indicate which tokens contributed to success or failure.

If \epsilon is very small (e.g., 10^{-8}), floating-point rounding errors or subnormal numbers cause \hat{A}_i to produce massive, erratic gradients. Conversely, if \epsilon is moderate, each completion receives \hat{A}_i = \frac{0}{\epsilon} = 0, which is mathematically correct but prone to floating-point drift.

More dangerously, if completions achieve identical partial format rewards (r_i = 0.1) without correct answers (r_{\text{acc}} = 0), calculating unconstrained advantages can result in updating policy weights toward repetitive thinking patterns.

The Zero-Variance Clamp Guardrail

To maintain strict mathematical stability across distributed ranks, the advantage calculation must enforce an explicit zero-variance guard:

python(16 lines)
1def compute_group_advantage(
2 rewards: list[float],
3 epsilon: float = -8
4) -> list[float]:
5 G = len(rewards)
6 mean_r = sum(rewards) / G
7 variance = sum((r - mean_r) ** 2 for r in rewards) / G
8 std_dev = variance ** 0.5
9
10 # Zero-Variance Clamp: If all completions in group performed identically,
11 # zero out all advantages to eliminate uninformative gradient noise.
12 if std_dev < -8:
13 return [0.0] * G
14
15 return [(r - mean_r) / (std_dev + epsilon) for r in rewards]

When \sigma_r < 10^{-8}, setting \hat{A}_i = 0.0 forces the clipped surrogate policy gradient \nabla_\theta \mathcal{J}(\theta) to zero for that group prompt, effectively bypassing parameter updates on non-informative rollouts.


Decisions

| Decision | Required evidence | Review trigger | |---|---|---| | Select GRPO over PPO for reasoning post-training | Verification that target domain admits deterministic rule-based verifiers (math, code, structured JSON). | Migrating post-training pipelines from conversational chat to mathematical/algorithmic reasoning. | | Set Group Size to G = 8 for dense models and G = 16 for MoE | Empirical validation that group variance \sigma_r^2 > 0 on at least 70% of training prompts. | Observational telemetry indicating > 40\% zero-variance clamp engagement. | | Enforce Schulman k_3 KL penalty with coefficient \beta \in [0.01, 0.05] | KL divergence monitoring showing strictly non-negative step metrics and absence of policy collapse. | Mean sequence length exploding by > 2.5\times baseline without benchmark accuracy gains. | | Isolate Code/Math Verifiers in gVisor / Firecracker microVMs | Benchmark execution latency under 250ms per test suite with complete network namespace isolation. | Deploying unvetted multi-turn code synthesis RL pipelines to distributed training nodes. |


Alternatives and trade-offs

text(11 lines)
1+-------------------+---------------------+--------------------+--------------------+
2| Attribute | Actor-Critic PPO | Critic-Free GRPO | DPO / KTO |
3+-------------------+---------------------+--------------------+--------------------+
4| Architecture | 4 Models (Heavy) | 2 Models (Lean) | 2 Models (Offline) |
5| Memory Footprint | + + + | + (50% save)| + |
6| Value Network | Required (V_phi) | None (Critic-Free) | None |
7| Exploration Type | Online On-Policy | Online On-Policy | Offline Fixed Pairs|
8| Reward Signal | Neural RM (Hackable)| Deterministic Rules| Static Preferences |
9| Self-Correction | Weak | Emergent in CoT | Not Learned |
10+-------------------+---------------------+--------------------+--------------------+
  • Direct Preference Optimization (DPO): While DPO is computationally lightweight (requiring no online sampling during parameter updates), it operates over static, pre-collected preference pairs. DPO cannot discover novel self-correction strategies or multi-step reasoning trajectories that were absent from the offline dataset.
  • Actor-Critic PPO: Remains mandatory for subjective tasks (creative writing, general conversational tone, nuanced persona alignment) where deterministic verification is impossible and a learned reward model is required.
  • Group Relative Policy Optimization (GRPO): Optimal for objective tasks with deterministic ground-truth verification. Eliminating the Critic cuts cluster VRAM requirements in half while group normalization provides robust contrastive signals for policy improvement.

Failure modes

  • The Zero-Variance Optimization Plateau: Occurs when training prompts are either too simple (100% group success) or impossibly complex (0% group success). With \sigma_r = 0, advantages evaluate to zero, causing gradient starvation and halting learning progression.
  • Runaway Chain-of-Thought Reflection Loops: If length penalties are insufficiently calibrated, the model discovers that generating circular reflective markers (e.g., "Wait, let me rethink this...") delays completion while preserving format rewards, resulting in maximum token limit saturation.
  • Language Divergence & Code Switching: During pure RLVR on English reasoning datasets, models initialized from multilingual checkpoints may spontaneously switch languages mid-trajectory if token probability distributions in another language happen to yield lower entropy. Mitigated by adding language consistency verification to r_{\text{format}}.
  • Format Reward Exploitation: The policy generates thousands of empty XML tags or injects multiple closing tags to fool naive regex parsers into assigning maximum format points without executing valid reasoning steps.

Operational checklist

  • [ ] All reward evaluation pipelines execute inside hardened, network-isolated sandboxes with execution timeouts under 500ms.
  • [ ] Group advantage normalization includes an explicit zero-variance guard (\sigma_r < 10^{-8} \implies \hat{A}_i = 0.0).
  • [ ] Schulman k_3 KL divergence formulation is validated to remain non-negative across all distributed training ranks.
  • [ ] Group rollout generation utilizes PagedAttention to eliminate KV-cache memory fragmentation during variable-length rollouts.
  • [ ] Training telemetry monitors the Zero-Variance Engagement Rate (\% \text{ of batches clamped}); if rate exceeds 35%, dynamic curriculum rebalancing is triggered.
  • [ ] Sequence length limits include safety margins to prevent distributed all-gather deadlocks when multiple completions hit max tokens.

Connected practice


Sources

  • deepseekmath-grpo-2024
  • deepseek-r1-paper