A Mental Model to Unify RL Losses
The diagram that compares PPO/GRPO, DAPO, GSPO, SAO, SAPO...
Inspired by the Feynman technique, I’ll explain some famous LLM RL losses as simply as I can—starting with PPO1 and GRPO2, then recent variants DAPO3, GSPO4, SAO5, and SAPO6. Writing them out helps me understand them better, and hopefully it helps you too.
💡 New to PPO or GRPO? These introductions are good places to start before reading this blog: PPO and GRPO.
Policy losses, or objectives, in LLM RL usually involve importance ratio. Define the token-level ratio as
where is the current policy and is the old policy. The old policy could be the snapshot of the model weights at the beginning of the gradient update, as in PPO, or the rollout policy in asynchronous training such as SAO5, lagged by more than one gradient update.
Let’s take PPO for example to see how is used in the RL. The full PPO objective1 is
To make it simpler, we can consider the token-level objective alone, since the full objective is an expectation of it over trajectories.
The min and clip operations in eq. (3) are where most writeups bury readers in notation. Split the objective by the sign of , then by where sits relative to the clip bounds, and the logic becomes straightforward.
Illustration of PPO
Case
This token is better than we thought, so we should increase its probability. Indeed,
Unclipped region ()
Increasing the better token probability increases the objective , exactly what we should do for a positive advantage .
Clipped region ().
In other words, it tells us don’t be too greedy if the current policy already strongly prefers the better token — we should stop assigning more importance to it.
Case
This token is worse than we thought, so we should decrease its probability. Indeed,
Unclipped region ()
Decreasing increases , as it should for a negative advantage.
Clipped region ().
Once the policy has already moved far enough against this token, further decreases are clipped out.
Figure 1 summarizes this behavior. On the vs. plane, a sampled token falls into one of four regions: in the red region, PPO/GRPO encourages the token by increasing its probability; in the blue region, it discourages the token; in the remaining regions, the gradient on is zero, so the token is effectively dropped from the parameter update.
Generalized RL objective
The SAPO paper6 gives us a useful way to compare these policy optimization methods. At the token level, write the surrogate objective as
Policy optimization methods, such as PPO/GRPO, DAPO, GSPO, SAO, and SAPO, differ mainly in their choice of the weight function . Taking the gradient of eq. (4),
All methods share the term . They only differ in the gate function that determines how much learning signal gets through.
Table 1 summarizes the weight function and gate function of some milestone policy optimization methods in LLM RL.
| Method | Insight | ||
|---|---|---|---|
| PPO (July 2017) / GRPO (February 2024) | Hard, asymmetric gating based on the sign of | ||
| DAPO (March 2025) | A higher upper bound leaves more room to increase useful low-probability tokens | ||
| GSPO (July 2025) | Same as DAPO at : |
One clip decision per response instead of per token | |
| SAPO (November 2025) | , if , else |
The gate decays smoothly instead of switching abruptly to zero; makes it decay faster for negative advantages | |
| SAO (July 2026) | Both signs are masked whenever the ratio leaves the trust region |
Figures 2–6 plot the gate and the learning signal on the plane for each method (Figure 6 uses the sequence pair ):
- PPO1 (Figure 2) clips the importance ratio : once an update has encouraged a good token or discouraged a bad token enough, the gate drops to zero and the token stops contributing. This keeps policy updates small and training stable.
- GRPO2 (Figure 2) adds a group rollout dimension to the objective. At the token level, its gate is identical to PPO’s.
- DAPO3 (Figure 3) raises the upper clipping threshold , so rare tokens—whose small denominator tends to produce large values—are less likely to be clipped prematurely. The extra strip above PPO’s boundary is exactly where the largest learning signals live.
- SAO5 (Figure 4) targets asynchronous training, where the current and rollout policies can drift farther apart than in PPO. can then deviate substantially from 1, so the gate behaves like a top-hat window, masking both signs whenever the ratio leaves the trust region.
- SAPO6 (Figure 5) replaces this hard boundary with a smooth gate. Instead of switching off a token’s contribution once crosses a threshold, SAPO lets the weight decay gradually as the ratio moves away from the preferred region.
- GSPO4 (Figure 6) is the same move at sequence level: swap for sequence level with —the geometric mean of the token ratios in eq. (1). Note that the shared factor changes as well: , i.e. every token in the response receives the same length-averaged weight. The gate is DAPO’s asymmetric two-wedge shape, but with GSPO’s clip range () stays within of 1, so the learning signal is essentially alone.
Takeaways
Using SAPO’s6 generalized form of objective with weight function , learning signals of different RL methods can be shown on the same plane.
- PPO / GRPO: hard token-level gate.
- DAPO: widens the gate for useful low-probability tokens.
- SAO: keeps only ratios inside a trust region when rollout policies become stale.
- SAPO: turns the hard gate into a smooth one.
- GSPO: moves the same idea from tokens to sequences.
Seen this way, these methods are providing different answers to the same question: which policy-gradient signals should we trust, and how much?
-
Schulman et al., Proximal Policy Optimization Algorithms, arXiv:1707.06347 (2017). ↩ ↩2 ↩3
-
Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, arXiv:2402.03300 (2024). ↩ ↩2
-
Yu et al., DAPO: An Open-Source LLM Reinforcement Learning System at Scale, arXiv:2503.14476 (2025). ↩ ↩2
-
Zheng et al., Group Sequence Policy Optimization, arXiv:2507.18071 (2025). ↩ ↩2
-
Hou et al., Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning, arXiv:2607.07508 (2026). ↩ ↩2 ↩3
-
Gao et al., Soft Adaptive Policy Optimization, arXiv:2511.20347 (2025). See also the Qwen Team blog post on SAPO. ↩ ↩2 ↩3 ↩4