Inspired by the Feynman technique, I’ll explain some famous LLM RL losses as simply as I can—starting with PPO1 and GRPO2, then recent variants DAPO3, GSPO4, SAO5, and SAPO6. Writing them out helps me understand them better, and hopefully it helps you too.

💡 New to PPO or GRPO? These introductions are good places to start before reading this blog: PPO and GRPO.

Policy losses, or objectives, in LLM RL usually involve importance ratio. Define the token-level ratio as

ρt=πθ(yt∣x,y<t)πold(yt∣x,y<t),(1)\rho_t = \frac{\pi_{\theta}(y_t \mid x, y_{<t})}{\pi_{\text{old}}(y_t \mid x, y_{<t})}, \tag{1}

where πθ\pi_\theta is the current policy and πold\pi_{\text{old}} is the old policy. The old policy could be the snapshot of the model weights at the beginning of the gradient update, as in PPO, or the rollout policy in asynchronous training such as SAO5, lagged by more than one gradient update.

Let’s take PPO for example to see how ρt\rho_t is used in the RL. The full PPO objective1 is

J(θ)=Ex∼D,  {yt}t=1T∼πold(⋅∣x)[∑t=1∣y∣min⁡(ρtAt,  clip⁡(ρt,1−ϵ,1+ϵ)At)].(2)J(\theta) = \mathbb{E}_{x \sim D,\; \{y_t\}_{t=1}^{T} \sim \pi_{\text{old}}(\cdot \mid x)} \left[ \sum_{t=1}^{|y|} \min\left( \rho_t A_t,\; \operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)A_t \right) \right]. \tag{2}

To make it simpler, we can consider the token-level objective alone, since the full objective is an expectation of it over trajectories.

Jt(θ)=min⁡(ρtAt,  clip⁡(ρt,1−ϵ,1+ϵ)At).(3)J_t(\theta) = \min\left( \rho_t A_t,\; \operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)A_t \right). \tag{3}

The min and clip operations in eq. (3) are where most writeups bury readers in notation. Split the objective by the sign of AtA_t, then by where ρt\rho_t sits relative to the clip bounds, and the logic becomes straightforward.

Illustration of PPO

Case At>0A_t > 0

This token is better than we thought, so we should increase its probability. Indeed,

Jt(θ)=min⁡(ρtAt, clip⁡(ρt,1−ϵ,1+ϵ)At)=min⁡(ρt, clip⁡(ρt,1−ϵ,1+ϵ)) At.J_t(\theta) = \min\bigl(\rho_t A_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)A_t\bigr) = \min\bigl(\rho_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\bigr)\,A_t. min⁡(ρt, clip⁡(ρt,1−ϵ,1+ϵ))={ρt,ρt≤1+ϵ,1+ϵ,ρt>1+ϵ.\min\bigl(\rho_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\bigr) = \begin{cases} \rho_t, & \rho_t \le 1+\epsilon, \\[4pt] 1+\epsilon, & \rho_t > 1+\epsilon. \end{cases}

Unclipped region (ρt<1+ϵ\rho_t < 1 + \epsilon)

Jt(θ)=ρtAt,∂Jt∂ρt=At>0J_t(\theta) = \rho_t A_t, \qquad \frac{\partial J_t}{\partial \rho_t}=A_t > 0

Increasing the better token probability ρt\rho_t increases the objective JtJ_t, exactly what we should do for a positive advantage AtA_t.

Clipped region (ρt>1+ϵ\rho_t > 1+\epsilon).

Jt(θ)=(1+ϵ)At,∂Jt∂ρt=0J_t(\theta) = (1 + \epsilon) A_t, \qquad \frac{\partial J_t}{\partial \rho_t}= 0

In other words, it tells us don’t be too greedy if the current policy πθ\pi_{\theta} already strongly prefers the better token — we should stop assigning more importance to it.

Case At<0A_t < 0

This token is worse than we thought, so we should decrease its probability. Indeed,

Jt(θ)=min⁡(ρtAt, clip⁡(ρt,1−ϵ,1+ϵ)At)=max⁡(ρt, clip⁡(ρt,1−ϵ,1+ϵ)) At.J_t(\theta) = \min\bigl(\rho_t A_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)A_t\bigr) = \max\bigl(\rho_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\bigr)\,A_t. max⁡(ρt, clip⁡(ρt,1−ϵ,1+ϵ))={1−ϵ,ρt<1−ϵ,ρt,ρt≥1−ϵ.\max\bigl(\rho_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\bigr) = \begin{cases} 1-\epsilon, & \rho_t < 1-\epsilon, \\[4pt] \rho_t, & \rho_t \ge 1-\epsilon. \end{cases}

Unclipped region (ρt≥1−ϵ\rho_t \ge 1-\epsilon)

Jt(θ)=ρtAt,∂Jt∂ρt=At<0J_t(\theta) = \rho_t A_t, \qquad \frac{\partial J_t}{\partial \rho_t}=A_t < 0

Decreasing ρt\rho_t increases JtJ_t, as it should for a negative advantage.

Clipped region (ρt<1−ϵ\rho_t < 1-\epsilon).

Jt(θ)=(1−ϵ)At,∂Jt∂ρt=0J_t(\theta) = (1 - \epsilon) A_t, \qquad \frac{\partial J_t}{\partial \rho_t}= 0

Once the policy has already moved far enough against this token, further decreases are clipped out.

Figure 1 summarizes this behavior. On the log⁡ρt\log\rho_t vs. AtA_t plane, a sampled token falls into one of four regions: in the red region, PPO/GRPO encourages the token by increasing its probability; in the blue region, it discourages the token; in the remaining regions, the gradient on ρt\rho_t is zero, so the token is effectively dropped from the parameter update.

PPO clip schematic on the advantage–log-ratio plane
Figure 1. Schematic diagram for PPO/GRPO on the log⁡ρt−At\log{\rho_t}-A_t plane. Red is where we increase ρt\rho_t, blue is where we decrease it, the rest are clipped out. The y-axis is log⁡ρt\log\rho_t so that the ratio's asymmetric range (0,∞)(0,\infty) becomes symmetric about 00. The default ϵ=0.2\epsilon = 0.2 is used.

Generalized RL objective

The SAPO paper6 gives us a useful way to compare these policy optimization methods. At the token level, write the surrogate objective as

Jt(θ)=f(ρt(θ);At) At.(4)J_t(\theta)=f(\rho_t(\theta);A_t)\,A_t. \tag{4}

Policy optimization methods, such as PPO/GRPO, DAPO, GSPO, SAO, and SAPO, differ mainly in their choice of the weight function ff. Taking the gradient of eq. (4),

∇θJt=Atf′(ρt;At) ∇θρt=f′(ρt;At)⏟method-specificρt At∇θlog⁡πθ(yt∣x,y<t)⏞policy gradient⏟shared.(5)\begin{aligned} \nabla_\theta J_t &= A_t f'(\rho_t;A_t)\,\nabla_\theta\rho_t \\[4pt] &= \underbrace{f'(\rho_t;A_t)}_{\text{method-specific}} \underbrace{ \rho_t\, \overbrace{A_t\nabla_\theta \log\pi_\theta(y_t\mid x,y_{<t})}^{\text{policy gradient}} }_{\text{shared}}. \end{aligned} \tag{5}

All methods share the term ρtAt∇θlog⁡πθ\rho_t A_t\nabla_\theta\log\pi_\theta. They only differ in the gate function f′(ρt;At)f'(\rho_t;A_t) that determines how much learning signal gets through.

Table 1 summarizes the weight function ff and gate function f′f' of some milestone policy optimization methods in LLM RL.

Method f(ρt;At)f(\rho_t;A_t) f′(ρt;At)f'(\rho_t;A_t) Insight
PPO (July 2017) / GRPO (February 2024) At>0:  min⁡(ρt,1+ϵ)A_t>0:\;\min(\rho_t,1+\epsilon)
At≤0:  max⁡(ρt,1−ϵ)A_t\leq0:\;\max(\rho_t,1-\epsilon)
1[At>0,  ρt<1+ϵ]\mathbf{1}\left[A_t>0,\;\rho_t<1+\epsilon\right]
+  1[At<0,  ρt>1−ϵ]+\;\mathbf{1}\left[A_t<0,\;\rho_t>1-\epsilon\right]
Hard, asymmetric gating based on the sign of AtA_t
DAPO (March 2025) At>0:  min⁡(ρt,1+ϵh)A_t>0:\;\min(\rho_t,1+\epsilon_h)
At≤0:  max⁡(ρt,1−ϵl)A_t\leq0:\;\max(\rho_t,1-\epsilon_l)
1[At>0,  ρt<1+ϵh]\mathbf{1}\left[A_t>0,\;\rho_t<1+\epsilon_h\right]
+  1[At<0,  ρt>1−ϵl]+\;\mathbf{1}\left[A_t<0,\;\rho_t>1-\epsilon_l\right]
A higher upper bound leaves more room to increase useful low-probability tokens
GSPO (July 2025) Same as DAPO at (ρs,As)(\rho_s,A_s):
As>0:  min⁡(ρs,1+ϵh)A_s>0:\;\min(\rho_s,1+\epsilon_h)
As≤0:  max⁡(ρs,1−ϵl)A_s\leq0:\;\max(\rho_s,1-\epsilon_l)
1[As>0,  ρs<1+ϵh]\mathbf{1}\left[A_s>0,\;\rho_s<1+\epsilon_h\right]
+  1[As<0,  ρs>1−ϵl]+\;\mathbf{1}\left[A_s<0,\;\rho_s>1-\epsilon_l\right]
One clip decision per response instead of per token
SAPO (November 2025) 4τtσ(τt(ρt−1))\dfrac{4}{\tau_t}\sigma\left(\tau_t(\rho_t-1)\right),
τt=τpos\tau_t=\tau_{\text{pos}} if At>0A_t>0, else τneg\tau_{\text{neg}}
4σ(τt(ρt−1))(1−σ(τt(ρt−1)))4\sigma\left(\tau_t(\rho_t-1)\right)\left(1-\sigma\left(\tau_t(\rho_t-1)\right)\right)
=sech2(τt(ρt−1)2)=\mathrm{sech}^2\left(\tfrac{\tau_t(\rho_t-1)}{2}\right)
The gate decays smoothly instead of switching abruptly to zero; τneg>τpos\tau_{\text{neg}}>\tau_{\text{pos}} makes it decay faster for negative advantages
SAO (July 2026) f~(ρt)=clip⁡(ρt,1−ϵl,1+ϵh)\widetilde f(\rho_t)=\operatorname{clip}(\rho_t,1-\epsilon_l,1+\epsilon_h) 1[1−ϵl<ρt<1+ϵh]\mathbf{1}\left[1-\epsilon_l<\rho_t<1+\epsilon_h\right] Both signs are masked whenever the ratio leaves the trust region
Table 1. Weight function ff and gate function f′f' for milestone policy optimization methods in LLM RL.

Figures 2–6 plot the gate f′(ρt;At)f'(\rho_t;A_t) and the learning signal f′(ρt;At) ρtAtf'(\rho_t;A_t)\,\rho_t A_t on the (At,log⁡ρt)(A_t,\log\rho_t) plane for each method (Figure 6 uses the sequence pair (As,ρs)(A_s,\rho_s)):

  • PPO1 (Figure 2) clips the importance ratio ρt\rho_t: once an update has encouraged a good token or discouraged a bad token enough, the gate drops to zero and the token stops contributing. This keeps policy updates small and training stable.
  • GRPO2 (Figure 2) adds a group rollout dimension GG to the objective. At the token level, its gate is identical to PPO’s.
  • DAPO3 (Figure 3) raises the upper clipping threshold ϵh\epsilon_h, so rare tokens—whose small denominator tends to produce large ρt\rho_t values—are less likely to be clipped prematurely. The extra strip above PPO’s boundary is exactly where the largest learning signals live.
  • SAO5 (Figure 4) targets asynchronous training, where the current and rollout policies can drift farther apart than in PPO. ρt\rho_t can then deviate substantially from 1, so the gate behaves like a top-hat window, masking both signs whenever the ratio leaves the trust region.
  • SAPO6 (Figure 5) replaces this hard boundary with a smooth gate. Instead of switching off a token’s contribution once ρt\rho_t crosses a threshold, SAPO lets the weight decay gradually as the ratio moves away from the preferred region.
  • GSPO4 (Figure 6) is the same move at sequence level: swap (At,ρt)(A_t,\rho_t) for sequence level (As,ρs)(A_s,\rho_s) with ρs=(πθ(y∣x)/πold(y∣x))1/∣y∣\rho_s=(\pi_\theta(y\mid x)/\pi_{\mathrm{old}}(y\mid x))^{1/\lvert y\rvert}—the geometric mean of the token ratios in eq. (1). Note that the shared factor changes as well: ∇θρs=ρs⋅1∣y∣∑t∇θlog⁡πθ(yt∣x,y<t)\nabla_\theta\rho_s=\rho_s\cdot\frac{1}{\lvert y\rvert}\sum_t\nabla_\theta\log\pi_\theta(y_t\mid x,y_{<t}), i.e. every token in the response receives the same length-averaged weight. The gate is DAPO’s asymmetric two-wedge shape, but with GSPO’s clip range (ϵ∼3–4×10−4\epsilon\sim 3\text{–}4\times10^{-4}) ρs\rho_s stays within 10−310^{-3} of 1, so the learning signal is essentially AsA_s alone.
PPO gate and learning-signal heatmaps on the advantage–log-ratio plane
Figure 2. PPO/GRPO. Left: gate f′f'. Right: learning signal f′ρtAtf'\rho_t A_t. Active regions are two quadrant wedges; learning signal grows with ρt\rho_t in the surviving corners.
DAPO gate and learning-signal heatmaps on the advantage–log-ratio plane
Figure 3. DAPO. Same binary gate as Figure 2, with a wider upper wedge.
SAO gate and learning-signal heatmaps on the advantage–log-ratio plane
Figure 4. SAO. Top-hat gate.
SAPO gate and learning-signal heatmaps on the advantage–log-ratio plane
Figure 5. SAPO. Smooth gate (dashed lines: f′f' iso-levels). Right: over this crop, ρt\rho_t growth and gate decay nearly cancel.
GSPO gradient-weight heatmap on the sequence-advantage / log-sequence-ratio plane
Figure 6. GSPO on (As,log⁡ρs)(A_s,\log\rho_s). Same wedge shape as Figure 3, but the axis spans ±10−3\pm10^{-3} because GSPO's clip range is three orders of magnitude tighter.

Takeaways

Using SAPO’s6 generalized form of objective JJ with weight function f(ρt;At)f(\rho_t; A_t), learning signals of different RL methods can be shown on the same (At,log⁡ρt)(A_t,\log\rho_t) plane.

  • PPO / GRPO: hard token-level gate.
  • DAPO: widens the gate for useful low-probability tokens.
  • SAO: keeps only ratios inside a trust region when rollout policies become stale.
  • SAPO: turns the hard gate into a smooth one.
  • GSPO: moves the same idea from tokens to sequences.

Seen this way, these methods are providing different answers to the same question: which policy-gradient signals should we trust, and how much?

  1. Schulman et al., Proximal Policy Optimization Algorithms, arXiv:1707.06347 (2017). ↩ ↩2 ↩3

  2. Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, arXiv:2402.03300 (2024). ↩ ↩2

  3. Yu et al., DAPO: An Open-Source LLM Reinforcement Learning System at Scale, arXiv:2503.14476 (2025). ↩ ↩2

  4. Zheng et al., Group Sequence Policy Optimization, arXiv:2507.18071 (2025). ↩ ↩2

  5. Hou et al., Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning, arXiv:2607.07508 (2026). ↩ ↩2 ↩3

  6. Gao et al., Soft Adaptive Policy Optimization, arXiv:2511.20347 (2025). See also the Qwen Team blog post on SAPO. ↩ ↩2 ↩3 ↩4