Group Relative Policy Optimization (GRPO): Mathematical Foundations, Value-Free Advantage Estimation, Group Normalization Dynamics, and Scalable Reasoning RL

Post-training reinforcement learning (RL) has become the core driver of complex reasoning capabilities in frontier language models. While early alignment workflows focused on conversational preference modeling via Proximal Policy Optimization (PPO) or Direct Preference Optimization (DPO), scaling reinforcement learning to multi-step mathematical derivation and code generation revealed structural inefficiencies in classical Actor-Critic architectures. The primary operational constraint in tradit

9 min
Group Relative Policy Optimization (GRPO): Mathematical Foundations, Value-Free Advantage Estimation, Group Normalization Dynamics, and Scalable Reasoning RL

Post-training reinforcement learning (RL) has become the core driver of complex reasoning capabilities in frontier language models. While early alignment workflows focused on conversational preference modeling via Proximal Policy Optimization (PPO) or Direct Preference Optimization (DPO), scaling reinforcement learning to multi-step mathematical derivation and code generation revealed structural inefficiencies in classical Actor-Critic architectures.

The primary operational constraint in traditional Actor-Critic reinforcement learning is the value network (the Critic). Training a separate parametric model of equivalent parameter scale to the policy model doubles the active parameter footprint in GPU memory and introduces numerical instability during value approximation on long autoregressive sequences.

Group Relative Policy Optimization (GRPO), introduced by Shao et al. in DeepSeekMath and later scaled in DeepSeek-R1, resolves this bottleneck by removing the critic network entirely. Instead of estimating absolute state values through parametric regression, GRPO samples a cohort of G candidate outputs for every prompt and derives baseline statistics directly from the empirical reward distribution of the group.

+-----------------------------------------------------------------------------+
|               ACTOR-CRITIC (PPO) vs. GROUP RELATIVE (GRPO)                  |
+-----------------------------------------------------------------------------+
| PPO Architecture:                                                           |
|                                                                             |
| Prompt q ---> [ Actor Policy \pi_\theta ] ---------> Single Completion o   |
|                      |                                       |              |
|                      v                                       v              |
|               [ Critic Value V_\psi ]              [ Reward Model r_\phi ]  |
|                      |                                       |              |
|                      +-------------> GAE A_t <---------------+              |
|                                                                             |
| Memory Footprint: Policy (\theta) + Critic (\psi) + Ref (\theta_ref) + RM   |
+-----------------------------------------------------------------------------+
| GRPO Architecture:                                                          |
|                                                                             |
| Prompt q ---> [ Actor Policy \pi_\theta ] ---> { o_1, o_2, ..., o_G }       |
|                                                       |                     |
|                                                       v                     |
|                                            [ Verifier / Reward r(o_i) ]     |
|                                                       |                     |
|                                                       v                     |
|               Advantage A_i = (r_i - Mean({r})) / (Std({r}) + \epsilon)     |
|                                                                             |
| Memory Footprint: Policy (\theta) + Ref (\theta_ref) (Critic Eliminated)    |
+-----------------------------------------------------------------------------+

The Actor-Critic Bottleneck in Language Model RL

In standard PPO implementations for language models (Ouyang et al., 2022), the training pipeline maintains four separate neural network models concurrently:

  • Actor Network (πθ\pi_\theta): The active autoregressive language model generating text tokens and receiving gradient updates.
  • Critic Network (VψV_\psi): A value function model predicting expected cumulative returns from the current token state.
  • Reference Policy (πref\pi_{\text{ref}}): A frozen snapshot of the initial supervised fine-tuned (SFT) model used to compute Kullback-Leibler (KL) divergence penalties.
  • Reward Model (rϕr_\phi): A preference scoring model evaluating generated sequences.
Actor-Critic PPO vs Group Relative Policy Optimization Architecture

Generalized Advantage Estimation (GAE) and Memory Pressure

PPO updates policy parameters by maximizing a clipped surrogate objective weighted by token-level advantage estimates AtA_t:

JPPO(θ)=EqP(Q),oπθold(Oq)[1ot=1omin(πθ(otq,o<t)πθold(otq,o<t)At,clip(πθ(otq,o<t)πθold(otq,o<t),1ε,1+ε)At)]\mathcal{J}_{\text{PPO}}(\theta) = \mathbb{E}_{q \sim P(Q), o \sim \pi_{\theta_{\text{old}}}(O|q)} \left[ \frac{1}{|o|} \sum_{t=1}^{|o|} \min \left( \frac{\pi_\theta(o_t | q, o_{<t})}{\pi_{\theta_{\text{old}}}(o_t | q, o_{<t})} A_t, \text{clip}\left(\frac{\pi_\theta(o_t | q, o_{<t})}{\pi_{\theta_{\text{old}}}(o_t | q, o_{<t})}, 1-\varepsilon, 1+\varepsilon\right) A_t \right) \right]

In PPO, AtA_t is computed via Generalized Advantage Estimation (Schulman et al., 2015):

δtV=rt+γVψ(st+1)Vψ(st)\delta_t^V = r_t + \gamma V_\psi(s_{t+1}) - V_\psi(s_t)

AtGAE(γ,λ)=l=0(γλ)lδt+lVA_t^{\text{GAE}(\gamma, \lambda)} = \sum_{l=0}^{\infty} (\gamma \lambda)^l \delta_{t+l}^V

This framework introduces two severe bottlenecks when applied to long-context reasoning models:

  1. VRAM Footprint: The Critic VψV_\psi requires a comparable parameter scale to the Actor to prevent representation collapse on complex mathematical tasks. Storing optimizer states (AdamW first and second moments), activations, and parameters for both πθ\pi_\theta and VψV_\psi consumes roughly twice the training memory of pure supervised fine-tuning.
  2. Credit Assignment Sparsity: In reasoning tasks (mathematics, software engineering, theorem proving), reward signals are typically terminal and binary: the final answer is either correct (r=1r=1) or incorrect (r=0r=0). Learning an accurate token-by-token value estimate Vψ(st)V_\psi(s_t) across thousands of intermediate reasoning tokens creates high gradient variance and value drift.

Mathematical Formulation of GRPO

Group Relative Policy Optimization avoids training a parametric value function VψV_\psi. Instead, for each query qP(Q)q \sim P(Q), the policy πθold\pi_{\theta_{\text{old}}} samples a discrete cohort of GG distinct candidate completions:

{o1,o2,,oG}πθold(Oq)\{o_1, o_2, \dots, o_G\} \sim \pi_{\theta_{\text{old}}}(O | q)

Each completion oio_i is evaluated by a scoring function (either a neural reward model or a deterministic rule-based verifier), yielding a scalar reward vector r={r1,r2,,rG}\mathbf{r} = \{r_1, r_2, \dots, r_G\}.

Group Relative Advantage Estimation

The advantage A^i\hat{A}_i for each response oio_i is calculated by standardizing scalar rewards across the sampled group:

μr=1Gj=1Grj,σr=1Gj=1G(rjμr)2+ϵ\mu_{\mathbf{r}} = \frac{1}{G} \sum_{j=1}^G r_j, \quad \sigma_{\mathbf{r}} = \sqrt{\frac{1}{G} \sum_{j=1}^G (r_j - \mu_{\mathbf{r}})^2 + \epsilon}

A^i=riμrσr\hat{A}_i = \frac{r_i - \mu_{\mathbf{r}}}{\sigma_{\mathbf{r}}}

Where ϵ\epsilon is a small constant (such as 10810^{-8}) preventing division by zero when all sampled completions receive identical scores.

Example Reward Normalization within Group (G = 4):
Prompt q: "Solve for x: 3x + 12 = 27"

Completion o_1: Correct derivation, x = 5      -> Reward r_1 = 1.0
Completion o_2: Arithmetic error,   x = 7      -> Reward r_2 = 0.0
Completion o_3: Correct derivation, x = 5      -> Reward r_3 = 1.0
Completion o_4: Hallucinated steps, x = -3     -> Reward r_4 = 0.0

Group Mean   \mu_r = (1.0 + 0.0 + 1.0 + 0.0) / 4 = 0.50
Group StdDev \sigma_r = 0.50

Normalized Advantages:
o_1: (1.0 - 0.50) / 0.50 = +1.00  (Reinforced)
o_2: (0.0 - 0.50) / 0.50 = -1.00  (Penalized)
o_3: (1.0 - 0.50) / 0.50 = +1.00  (Reinforced)
o_4: (0.0 - 0.50) / 0.50 = -1.00  (Penalized)

Outcome vs. Process Advantage Assignment

Depending on the supervision granularity, token-level advantages A^i,t\hat{A}_{i,t} are assigned under two regimes:

  • Outcome Supervision (Terminal Reward): When scoring only the final output correctness, the standardized reward A^i\hat{A}_i is broadcast uniformly across all tokens in completion oio_i:

A^i,t=A^i=riμrσr,t[1,oi]\hat{A}_{i,t} = \hat{A}_i = \frac{r_i - \mu_{\mathbf{r}}}{\sigma_{\mathbf{r}}}, \quad \forall t \in [1, |o_i|]

  • Process Supervision (Step-Level Reward): When a process reward model evaluates step-by-step reasoning steps j[1,Ki]j \in [1, K_i], each step ending at token index index(j)\text{index}(j) receives a step reward riindex(j)r_i^{\text{index}(j)}. Rewards across all steps in the group are normalized to r~iindex(j)\widetilde{r}_i^{\text{index}(j)}, and token advantages are calculated via future cumulative sum:

A^i,t=index(j)tr~iindex(j)\hat{A}_{i,t} = \sum_{\text{index}(j) \ge t} \widetilde{r}_i^{\text{index}(j)}


Optimization Objective and Unbiased KL Regularization

The full optimization objective of GRPO optimizes the policy parameters θ\theta over prompt batches:

JGRPO(θ)=EqP(Q),{oi}i=1Gπθold(Oq)[1Gi=1G1oit=1oi(min(ρi,tA^i,t,clip(ρi,t,1ε,1+ε)A^i,t)βDKL[πθπref])]\mathcal{J}_{\text{GRPO}}(\theta) = \mathbb{E}_{q \sim P(Q), \{o_i\}_{i=1}^G \sim \pi_{\theta_{\text{old}}}(O|q)} \left[ \frac{1}{G} \sum_{i=1}^G \frac{1}{|o_i|} \sum_{t=1}^{|o_i|} \left( \min \left( \rho_{i,t} \hat{A}_{i,t}, \text{clip}(\rho_{i,t}, 1-\varepsilon, 1+\varepsilon) \hat{A}_{i,t} \right) - \beta \mathbb{D}_{\text{KL}}\left[\pi_\theta \parallel \pi_{\text{ref}}\right] \right) \right]

Where the per-token importance sampling probability ratio is defined as:

ρi,t=πθ(oi,tq,oi,<t)πθold(oi,tq,oi,<t)\rho_{i,t} = \frac{\pi_\theta(o_{i,t} \mid q, o_{i,<t})}{\pi_{\theta_{\text{old}}}(o_{i,t} \mid q, o_{i,<t})}

Direct Loss-Level KL Divergence

Unlike PPO, which typically injects a per-token KL penalty directly into the scalar reward formulation (rt=r(st,at)βlogπθπrefr_t = r(s_t, a_t) - \beta \log \frac{\pi_\theta}{\pi_{\text{ref}}}), GRPO incorporates the KL divergence directly into the loss function.

DeepSeekMath utilizes the non-negative unbiased KL divergence estimator derived by Schulman (2020):

DKL[πθπref]=πref(oi,tq,oi,<t)πθ(oi,tq,oi,<t)logπref(oi,tq,oi,<t)πθ(oi,tq,oi,<t)1\mathbb{D}_{\text{KL}}\left[\pi_\theta \parallel \pi_{\text{ref}}\right] = \frac{\pi_{\text{ref}}(o_{i,t} \mid q, o_{i,<t})}{\pi_\theta(o_{i,t} \mid q, o_{i,<t})} - \log \frac{\pi_{\text{ref}}(o_{i,t} \mid q, o_{i,<t})}{\pi_\theta(o_{i,t} \mid q, o_{i,<t})} - 1

This formulation guarantees:

  1. Non-Negativity: For any positive probability ratio x=πrefπθx = \frac{\pi_{\text{ref}}}{\pi_\theta}, the expression xlogx10x - \log x - 1 \ge 0, with equality holding if and only if x=1x = 1.
  2. Stable Variance: It avoids large negative penalty spikes that can destabilize policy updates when sampled sequences deviate slightly from the reference policy.

Comparison: Unified Reinforcement Learning Paradigm

In the DeepSeekMath analysis, post-training methods are unified under a generalized gradient framework:

θJA(θ)=E(q,o)D[1ot=1oGCA(q,o,t,πrf)θlogπθ(otq,o<t)]\nabla_\theta \mathcal{J}_{\mathcal{A}}(\theta) = \mathbb{E}_{(q, o) \sim \mathcal{D}} \left[ \frac{1}{|o|} \sum_{t=1}^{|o|} GC_{\mathcal{A}}(q, o, t, \pi_{\text{rf}}) \nabla_\theta \log \pi_\theta(o_t \mid q, o_{<t}) \right]

Where GCAGC_{\mathcal{A}} represents the gradient coefficient determining the magnitude and sign of the parameter update for each token.

+-----------------------------------------------------------------------------+
|                POST-TRAINING REINFORCEMENT LEARNING PARADIGMS               |
+-----------------------------------------------------------------------------+
| Algorithm: SFT                                                              |
| Sampling:  Offline (Static supervised dataset)                              |
| Gradient:  GC = 1.0 (Constant positive reinforcement)                       |
| VRAM Load: 1 Model (Policy \pi_\theta)                                      |
+-----------------------------------------------------------------------------+
| Algorithm: Rejection Sampling Fine-Tuning (RFT)                             |
| Sampling:  Offline (Sampled from \pi_sft, filtered on binary correctness)   |
| Gradient:  GC = 1.0 (Positive only for correct answers, zero otherwise)     |
| VRAM Load: 1 Model (Policy \pi_\theta)                                      |
+-----------------------------------------------------------------------------+
| Algorithm: Online Rejection Sampling (Online RFT)                           |
| Sampling:  Online (Sampled dynamically from active policy \pi_\theta)       |
| Gradient:  GC = 1.0 (Positive only for correct answers, zero otherwise)     |
| VRAM Load: 1 Model (Policy \pi_\theta)                                      |
+-----------------------------------------------------------------------------+
| Algorithm: Direct Preference Optimization (DPO)                             |
| Sampling:  Offline (Static paired completions (o+, o-) from \pi_sft)        |
| Gradient:  Implicit log-ratio preference margin                             |
| VRAM Load: 2 Models (Policy \pi_\theta + Ref \pi_ref)                       |
+-----------------------------------------------------------------------------+
| Algorithm: Proximal Policy Optimization (PPO)                               |
| Sampling:  Online (Single trajectory per prompt)                            |
| Gradient:  GC = Advantage A_t (Estimated via parametric Critic V_\psi)      |
| VRAM Load: 4 Models (\pi_\theta + \pi_old + Critic V_\psi + Ref \pi_ref)   |
+-----------------------------------------------------------------------------+
| Algorithm: Group Relative Policy Optimization (GRPO)                        |
| Sampling:  Online (Cohort of G completions per prompt)                      |
| Gradient:  GC = Normalized group advantage (Bidirectional: +/-)             |
| VRAM Load: 2 Models (Policy \pi_\theta + Ref \pi_ref)                       |
+-----------------------------------------------------------------------------+

Why GRPO Outperforms Online Rejection Sampling (Online RFT)

Online RFT only performs positive reinforcement on valid completions (GC=+1GC = +1) and drops failed attempts. In contrast, GRPO applies bidirectional updates:

  • High-scoring trajectories receive positive gradient coefficients (A^i>0\hat{A}_i > 0).
  • Low-scoring or incorrect trajectories within the same prompt cohort receive negative gradient coefficients (A^i<0\hat{A}_i < 0), actively pushing probability mass away from degenerate reasoning paths.
  • The magnitude of the coefficient scales continuously with the relative quality of the reasoning chain.

Memory Economics and Scaling Dynamics

By discarding the Critic network VψV_\psi, GRPO dramatically reduces per-node GPU memory requirements during distributed reinforcement learning.

+-----------------------------------------------------------------------------+
|               ESTIMATED MEMORY FOOTPRINT (70B PARAMETER MODEL)              |
+-----------------------------------------------------------------------------+
| Component                   PPO (16-bit + AdamW)      GRPO (16-bit + AdamW) |
+-----------------------------------------------------------------------------+
| Actor Weights (70B FP16)    140 GB                    140 GB                |
| Actor Optimizer (AdamW)     560 GB (FP32 states)      560 GB (FP32 states)  |
| Critic Weights (70B FP16)   140 GB                    0 GB (Eliminated)     |
| Critic Optimizer (AdamW)    560 GB                    0 GB (Eliminated)     |
| Reference Model (70B FP16)  140 GB                    140 GB                |
| Reward Model (70B FP16)     140 GB                    0-140 GB (0 if rule)  |
+-----------------------------------------------------------------------------+
| Total Parameter/Opt VRAM:   ~1,680 GB                 ~840 GB (-50% VRAM)   |
+-----------------------------------------------------------------------------+

When coupled with rule-based verifiers (such as a Python execution sandbox for code or math syntax parse trees for symbolic answers), the reward model footprint is eliminated as well, enabling full post-training RL on consumer or mid-scale compute clusters.


PyTorch Reference Implementation

The following self-contained PyTorch module illustrates group advantage computation, loss clipping, and Schulman KL divergence calculation in GRPO:

import torch
import torch.nn as nn
import torch.nn.functional as F

class GRPOTrainer:
    def __init__(
        self,
        clip_eps: float = 0.2,
        kl_beta: float = 0.04,
        eps: float = 1e-8
    ):
        self.clip_eps = clip_eps
        self.kl_beta = kl_beta
        self.eps = eps

    def compute_group_advantages(self, rewards: torch.Tensor) -> torch.Tensor:
        """
        Compute normalized group advantages for GRPO.
        Args:
            rewards: Tensor of shape (batch_size, group_size)
        Returns:
            advantages: Tensor of shape (batch_size, group_size)
        """
        mean = rewards.mean(dim=-1, keepdim=True)
        std = rewards.std(dim=-1, keepdim=True)
        advantages = (rewards - mean) / (std + self.eps)
        return advantages

    def compute_loss(
        self,
        log_probs: torch.Tensor,         # Shape: (B * G, seq_len)
        old_log_probs: torch.Tensor,     # Shape: (B * G, seq_len)
        ref_log_probs: torch.Tensor,     # Shape: (B * G, seq_len)
        advantages: torch.Tensor,        # Shape: (B * G, 1)
        completion_mask: torch.Tensor    # Shape: (B * G, seq_len)
    ) -> torch.Tensor:
        """
        Compute the GRPO surrogate objective with Schulman KL divergence.
        """
        # Probability ratio: exp(log \pi_\theta - log \pi_{\theta_old})
        ratio = torch.exp(log_probs - old_log_probs)

        # Clipped surrogate objective
        surr1 = ratio * advantages
        surr2 = torch.clamp(ratio, 1.0 - self.clip_eps, 1.0 + self.clip_eps) * advantages
        policy_loss = -torch.min(surr1, surr2)

        # Unbiased Schulman KL divergence: (p_ref / p_theta) - log(p_ref / p_theta) - 1
        # In log space: ratio_ref = exp(ref_log_probs - log_probs)
        ratio_ref = torch.exp(ref_log_probs - log_probs)
        kl_div = ratio_ref - (ref_log_probs - log_probs) - 1.0

        # Total token loss
        token_loss = policy_loss + self.kl_beta * kl_div

        # Mask padding tokens and average over sequence length and batch
        masked_loss = (token_loss * completion_mask).sum(dim=-1) / completion_mask.sum(dim=-1).clamp(min=1.0)
        return masked_loss.mean()


if __name__ == "__main__":
    batch_size = 2
    group_size = 4
    seq_len = 16

    trainer = GRPOTrainer(clip_eps=0.2, kl_beta=0.04)

    # Simulated rewards from verifier for 2 queries, 4 outputs each
    rewards = torch.tensor([
        [1.0, 0.0, 1.0, 0.0],
        [0.0, 0.0, 1.0, 1.0]
    ])

    # Compute advantages: (2, 4) -> reshape to (8, 1)
    advantages = trainer.compute_group_advantages(rewards).view(-1, 1)

    # Simulated log probabilities
    log_probs = torch.randn(batch_size * group_size, seq_len)
    old_log_probs = log_probs.detach() + torch.randn_like(log_probs) * 0.05
    ref_log_probs = log_probs.detach() + torch.randn_like(log_probs) * 0.1
    completion_mask = torch.ones(batch_size * group_size, seq_len)

    loss = trainer.compute_loss(
        log_probs, old_log_probs, ref_log_probs, advantages, completion_mask
    )
    print(f"GRPO Surrogate Loss: {loss.item():.4f}")

Sources

Written by

More to read

  • Embedding Models and Rerankers in Production: Comparing Dense Bi-Encoders, ColBERT Late Interaction, Cross-Encoders, and Serving Architectures

    Production retrieval-augmented generation systems depend heavily on the retrieval pipeline that precedes generation. While frontier language models receive the bulk of engineering attention, the accuracy, latency, and operational cost of a RAG application are often determined by the interplay between embedding models and rerankers. Deploying search infrastructure requires navigating three distinct architectural paradigms: dense bi-encoders, multi-vector late interaction engines, and cross-encod

    1 min
  • Anthropic Opens 250,000 Claude Conversations to External Researchers Across Stanford, Oxford, and METR

    Anthropic's Societal Impacts team has released initial results from a research pilot that opened aggregate, real-world Claude conversation data to outside academic teams. The initiative partnered with researchers from Stanford University's Social and Language Technologies (SALT) Lab, the University of Oxford's Human Information Processing Lab, and the model evaluation non-profit METR. Each group conducted independent studies across a sample of approximately 250,000 conversations recorded on Clau

    1 min
  • Model Context Protocol (MCP) in Production AI Agents: Architecture, Transport Layers, Security Sandboxing, and Tool Federation

    Model Context Protocol (MCP) in Production AI Agents: Architecture, Transport Layers, Security Sandboxing, and Tool Federation The transition from standalone large language models to autonomous agentic systems has introduced an integration scaling problem. Early agent implementations relied on proprietary, ad hoc function-calling wrappers written specifically for each model provider or orchestration framework. Connecting $M$ distinct agent runtimes to $N$ enterprise data stores and developer to

    1 min