<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="3.9.5">Jekyll</generator><link href="https://wang-zhongwei.github.io/blog/feed.xml" rel="self" type="application/atom+xml" /><link href="https://wang-zhongwei.github.io/blog/" rel="alternate" type="text/html" hreflang="en" /><updated>2026-08-17T03:57:15+00:00</updated><id>https://wang-zhongwei.github.io/blog/feed.xml</id><title type="html">Near Criticality</title><subtitle>Notes by Zhongwei Wang on RL, AI, physics and complex systems</subtitle><author><name>Zhongwei Wang</name></author><entry><title type="html">A Mental Model to Unify RL Losses</title><link href="https://wang-zhongwei.github.io/blog/2026/08/14/mental-model-policy-optimizations.html" rel="alternate" type="text/html" title="A Mental Model to Unify RL Losses" /><published>2026-08-14T00:00:00+00:00</published><updated>2026-08-14T00:00:00+00:00</updated><id>https://wang-zhongwei.github.io/blog/2026/08/14/mental-model-policy-optimizations</id><content type="html" xml:base="https://wang-zhongwei.github.io/blog/2026/08/14/mental-model-policy-optimizations.html"><![CDATA[<p>Inspired by the <a href="https://en.wikipedia.org/wiki/Learning_by_teaching">Feynman technique</a>, I’ll explain some famous LLM RL losses as simply as I can—starting with PPO<sup id="fnref:ppo"><a href="#fn:ppo" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> and GRPO<sup id="fnref:grpo"><a href="#fn:grpo" class="footnote" rel="footnote" role="doc-noteref">2</a></sup>, then recent variants DAPO<sup id="fnref:dapo"><a href="#fn:dapo" class="footnote" rel="footnote" role="doc-noteref">3</a></sup>, GSPO<sup id="fnref:gspo"><a href="#fn:gspo" class="footnote" rel="footnote" role="doc-noteref">4</a></sup>, SAO<sup id="fnref:sao"><a href="#fn:sao" class="footnote" rel="footnote" role="doc-noteref">5</a></sup>, and SAPO<sup id="fnref:sapo"><a href="#fn:sapo" class="footnote" rel="footnote" role="doc-noteref">6</a></sup>. Writing them out helps me understand them better, and hopefully it helps you too.</p>

<blockquote>
  <p>💡 <strong>New to PPO or GRPO?</strong> These introductions are good places to start before reading this blog: <a href="https://huggingface.co/blog/deep-rl-ppo">PPO</a> and <a href="https://huggingface.co/blog/garg-aayush/derive-grpo-loss">GRPO</a>.</p>
</blockquote>

<p>Policy losses, or objectives, in LLM RL usually involve importance ratio. Define the token-level ratio as</p>

\[\rho_t = \frac{\pi_{\theta}(y_t \mid x, y_{&lt;t})}{\pi_{\text{old}}(y_t \mid x, y_{&lt;t})},
\tag{1}\]

<p>where \(\pi_\theta\) is the current policy and \(\pi_{\text{old}}\) is the old policy. The old policy could be the snapshot of the model weights at the beginning of the gradient update, as in PPO, or the rollout policy in asynchronous training such as SAO<sup id="fnref:sao:1"><a href="#fn:sao" class="footnote" rel="footnote" role="doc-noteref">5</a></sup>, lagged by more than one gradient update.</p>

<p>Let’s take PPO for example to see how \(\rho_t\) is used in the RL. The full PPO objective<sup id="fnref:ppo:1"><a href="#fn:ppo" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> is</p>

\[J(\theta)
=
\mathbb{E}_{x \sim D,\; \{y_t\}_{t=1}^{T} \sim \pi_{\text{old}}(\cdot \mid x)}
\left[
\sum_{t=1}^{|y|}
\min\left(
\rho_t A_t,\;
\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)A_t
\right)
\right].
\tag{2}\]

<p>To make it simpler, we can consider the token-level objective alone, since the full objective is an expectation of it over trajectories.</p>

\[J_t(\theta) = \min\left(
\rho_t A_t,\;
\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)A_t
\right).
\tag{3}\]

<p>The <code class="language-plaintext highlighter-rouge">min</code> and <code class="language-plaintext highlighter-rouge">clip</code> operations in eq. (3) are where most writeups bury readers in notation. Split the objective by the sign of \(A_t\), then by where \(\rho_t\) sits relative to the clip bounds, and the logic becomes straightforward.</p>

<h2 id="illustration-of-ppo">Illustration of PPO</h2>

<h3 id="case-a_t--0">Case \(A_t &gt; 0\)</h3>

<p>This token is better than we thought, so we should increase its probability. Indeed,</p>

\[J_t(\theta)
=
\min\bigl(\rho_t A_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)A_t\bigr)
=
\min\bigl(\rho_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\bigr)\,A_t.\]

\[\min\bigl(\rho_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\bigr)
=
\begin{cases}
\rho_t, &amp; \rho_t \le 1+\epsilon, \\[4pt]
1+\epsilon, &amp; \rho_t &gt; 1+\epsilon.
\end{cases}\]

<p><strong>Unclipped region (\(\rho_t &lt; 1 + \epsilon\))</strong></p>

\[J_t(\theta) = \rho_t A_t,
\qquad
\frac{\partial J_t}{\partial \rho_t}=A_t &gt; 0\]

<p>Increasing the better token probability \(\rho_t\) increases the objective \(J_t\), <em>exactly what we should</em> do for a positive advantage \(A_t\).</p>

<p><strong>Clipped region (\(\rho_t &gt; 1+\epsilon\)).</strong></p>

\[J_t(\theta) = (1 + \epsilon) A_t,
\qquad
\frac{\partial J_t}{\partial \rho_t}= 0\]

<p>In other words, it tells us <em>don’t be too greedy</em> if the current policy \(\pi_{\theta}\) already <em>strongly prefers</em> the better token — we should stop assigning more importance to it.</p>

<h3 id="case-a_t--0-1">Case \(A_t &lt; 0\)</h3>

<p>This token is worse than we thought, so we should decrease its probability. Indeed,</p>

\[J_t(\theta)
=
\min\bigl(\rho_t A_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)A_t\bigr)
=
\max\bigl(\rho_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\bigr)\,A_t.\]

\[\max\bigl(\rho_t,\,\operatorname{clip}(\rho_t, 1-\epsilon, 1+\epsilon)\bigr)
=
\begin{cases}
1-\epsilon, &amp; \rho_t &lt; 1-\epsilon, \\[4pt]
\rho_t, &amp; \rho_t \ge 1-\epsilon.
\end{cases}\]

<p><strong>Unclipped region (\(\rho_t \ge 1-\epsilon\))</strong></p>

\[J_t(\theta) = \rho_t A_t,
\qquad
\frac{\partial J_t}{\partial \rho_t}=A_t &lt; 0\]

<p>Decreasing \(\rho_t\) increases \(J_t\), as it should for a negative advantage.</p>

<p><strong>Clipped region (\(\rho_t &lt; 1-\epsilon\)).</strong></p>

\[J_t(\theta) = (1 - \epsilon) A_t,
\qquad
\frac{\partial J_t}{\partial \rho_t}= 0\]

<p>Once the policy has already moved far enough against this token, further decreases are clipped out.</p>

<p><a href="#figure-ppo-clip-schematic">Figure 1</a> summarizes this behavior. On the \(\log\rho_t\) vs. \(A_t\) plane, a sampled token falls into one of four regions: in the red region, PPO/GRPO encourages the token by increasing its probability; in the blue region, it discourages the token; in the remaining regions, the gradient on \(\rho_t\) is zero, so the token is effectively dropped from the parameter update.</p>

<figure id="figure-ppo-clip-schematic">
  <img src="/blog/assets/figures/mental-model-policy-optimizations/ppo-clip-schematic.png" alt="PPO clip schematic on the advantage–log-ratio plane" />
  <figcaption><strong>Figure 1.</strong> Schematic diagram for PPO/GRPO on the \(\log{\rho_t}-A_t\) plane. Red is where we increase \(\rho_t\), blue is where we decrease it, the rest are clipped out. The y-axis is \(\log\rho_t\) so that the ratio's asymmetric range \((0,\infty)\) becomes symmetric about \(0\). The default \(\epsilon = 0.2\) is used. </figcaption>
</figure>

<h2 id="generalized-rl-objective">Generalized RL objective</h2>

<p>The SAPO paper<sup id="fnref:sapo:1"><a href="#fn:sapo" class="footnote" rel="footnote" role="doc-noteref">6</a></sup> gives us a useful way to compare these policy optimization methods. At the token level, write the surrogate objective as</p>

\[J_t(\theta)=f(\rho_t(\theta);A_t)\,A_t.
\tag{4}\]

<p>Policy optimization methods, such as PPO/GRPO, DAPO, GSPO, SAO, and SAPO, differ mainly in their choice of the <strong>weight function</strong> \(f\). Taking the gradient of eq. (4),</p>

\[\begin{aligned}
\nabla_\theta J_t
&amp;=
A_t f'(\rho_t;A_t)\,\nabla_\theta\rho_t \\[4pt]
&amp;=
\underbrace{f'(\rho_t;A_t)}_{\text{method-specific}}
\underbrace{
\rho_t\,
\overbrace{A_t\nabla_\theta
\log\pi_\theta(y_t\mid x,y_{&lt;t})}^{\text{policy gradient}}
}_{\text{shared}}.
\end{aligned}
\tag{5}\]

<p>All methods share the term \(\rho_t A_t\nabla_\theta\log\pi_\theta\). They only differ in the <strong>gate function</strong> \(f'(\rho_t;A_t)\) that determines how much learning signal gets through.</p>

<p><a href="#table-1">Table 1</a> summarizes the weight function \(f\) and gate function \(f'\) of some milestone policy optimization methods in LLM RL.</p>

<figure id="table-1">

  <table>
    <thead>
      <tr>
        <th>Method</th>
        <th>\(f(\rho_t;A_t)\)</th>
        <th>\(f'(\rho_t;A_t)\)</th>
        <th>Insight</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td><a href="https://arxiv.org/abs/1707.06347">PPO</a> (July 2017) / <a href="https://arxiv.org/abs/2402.03300">GRPO</a> (February 2024)</td>
        <td>\(A_t&gt;0:\;\min(\rho_t,1+\epsilon)\)<br />\(A_t\leq0:\;\max(\rho_t,1-\epsilon)\)</td>
        <td>\(\mathbf{1}\left[A_t&gt;0,\;\rho_t&lt;1+\epsilon\right]\)<br />\(+\;\mathbf{1}\left[A_t&lt;0,\;\rho_t&gt;1-\epsilon\right]\)</td>
        <td>Hard, asymmetric gating based on the sign of \(A_t\)</td>
      </tr>
      <tr>
        <td><a href="https://arxiv.org/abs/2503.14476">DAPO</a> (March 2025)</td>
        <td>\(A_t&gt;0:\;\min(\rho_t,1+\epsilon_h)\)<br />\(A_t\leq0:\;\max(\rho_t,1-\epsilon_l)\)</td>
        <td>\(\mathbf{1}\left[A_t&gt;0,\;\rho_t&lt;1+\epsilon_h\right]\)<br />\(+\;\mathbf{1}\left[A_t&lt;0,\;\rho_t&gt;1-\epsilon_l\right]\)</td>
        <td>A higher upper bound leaves more room to increase useful low-probability tokens</td>
      </tr>
      <tr>
        <td><a href="https://arxiv.org/abs/2507.18071">GSPO</a> (July 2025)</td>
        <td>Same as DAPO at \((\rho_s,A_s)\):<br />\(A_s&gt;0:\;\min(\rho_s,1+\epsilon_h)\)<br />\(A_s\leq0:\;\max(\rho_s,1-\epsilon_l)\)</td>
        <td>\(\mathbf{1}\left[A_s&gt;0,\;\rho_s&lt;1+\epsilon_h\right]\)<br />\(+\;\mathbf{1}\left[A_s&lt;0,\;\rho_s&gt;1-\epsilon_l\right]\)</td>
        <td>One clip decision per response instead of per token</td>
      </tr>
      <tr>
        <td><a href="https://arxiv.org/abs/2511.20347">SAPO</a> (November 2025)</td>
        <td>\(\dfrac{4}{\tau_t}\sigma\left(\tau_t(\rho_t-1)\right)\),<br />\(\tau_t=\tau_{\text{pos}}\) if \(A_t&gt;0\), else \(\tau_{\text{neg}}\)</td>
        <td>\(4\sigma\left(\tau_t(\rho_t-1)\right)\left(1-\sigma\left(\tau_t(\rho_t-1)\right)\right)\)<br />\(=\mathrm{sech}^2\left(\tfrac{\tau_t(\rho_t-1)}{2}\right)\)</td>
        <td>The gate decays smoothly instead of switching abruptly to zero; \(\tau_{\text{neg}}&gt;\tau_{\text{pos}}\) makes it decay faster for negative advantages</td>
      </tr>
      <tr>
        <td><a href="https://arxiv.org/abs/2607.07508">SAO</a> (July 2026)</td>
        <td>\(\widetilde f(\rho_t)=\operatorname{clip}(\rho_t,1-\epsilon_l,1+\epsilon_h)\)</td>
        <td>\(\mathbf{1}\left[1-\epsilon_l&lt;\rho_t&lt;1+\epsilon_h\right]\)</td>
        <td>Both signs are masked whenever the ratio leaves the trust region</td>
      </tr>
    </tbody>
  </table>

  <figcaption><strong>Table 1.</strong> Weight function \(f\) and gate function \(f'\) for milestone policy optimization methods in LLM RL.</figcaption>
</figure>

<p>Figures 2–6 plot the gate \(f'(\rho_t;A_t)\) and the learning signal \(f'(\rho_t;A_t)\,\rho_t A_t\) on the \((A_t,\log\rho_t)\) plane for each method (Figure 6 uses the sequence pair \((A_s,\rho_s)\)):</p>

<ul>
  <li>PPO<sup id="fnref:ppo:2"><a href="#fn:ppo" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> (<a href="#figure-ppo-gradient-weight">Figure 2</a>) clips the importance ratio \(\rho_t\): once an update has encouraged a good token or discouraged a bad token enough, the gate drops to zero and the token stops contributing. This keeps policy updates small and training stable.</li>
  <li>GRPO<sup id="fnref:grpo:1"><a href="#fn:grpo" class="footnote" rel="footnote" role="doc-noteref">2</a></sup> (<a href="#figure-ppo-gradient-weight">Figure 2</a>) adds a group rollout dimension \(G\) to the objective. At the token level, its gate is identical to PPO’s.</li>
  <li>DAPO<sup id="fnref:dapo:1"><a href="#fn:dapo" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> (<a href="#figure-dapo-gradient-weight">Figure 3</a>) raises the upper clipping threshold \(\epsilon_h\), so rare tokens—whose small denominator tends to produce large \(\rho_t\) values—are less likely to be clipped prematurely. The extra strip above PPO’s boundary is exactly where the largest learning signals live.</li>
  <li>SAO<sup id="fnref:sao:2"><a href="#fn:sao" class="footnote" rel="footnote" role="doc-noteref">5</a></sup> (<a href="#figure-sao-gradient-weight">Figure 4</a>) targets asynchronous training, where the current and rollout policies can drift farther apart than in PPO. \(\rho_t\) can then deviate substantially from 1, so the gate behaves like a top-hat window, masking both signs whenever the ratio leaves the trust region.</li>
  <li>SAPO<sup id="fnref:sapo:2"><a href="#fn:sapo" class="footnote" rel="footnote" role="doc-noteref">6</a></sup> (<a href="#figure-sapo-gradient-weight">Figure 5</a>) replaces this hard boundary with a smooth gate. Instead of switching off a token’s contribution once \(\rho_t\) crosses a threshold, SAPO lets the weight decay gradually as the ratio moves away from the preferred region.</li>
  <li>GSPO<sup id="fnref:gspo:1"><a href="#fn:gspo" class="footnote" rel="footnote" role="doc-noteref">4</a></sup> (<a href="#figure-gspo-gradient-weight">Figure 6</a>) is the same move at sequence level: swap \((A_t,\rho_t)\) for sequence level \((A_s,\rho_s)\) with \(\rho_s=(\pi_\theta(y\mid x)/\pi_{\mathrm{old}}(y\mid x))^{1/\lvert y\rvert}\)—the geometric mean of the token ratios in eq. (1). Note that the shared factor changes as well: \(\nabla_\theta\rho_s=\rho_s\cdot\frac{1}{\lvert y\rvert}\sum_t\nabla_\theta\log\pi_\theta(y_t\mid x,y_{&lt;t})\), i.e. every token in the response receives the same length-averaged weight. The gate is DAPO’s asymmetric two-wedge shape, but with GSPO’s clip range (\(\epsilon\sim 3\text{–}4\times10^{-4}\)) \(\rho_s\) stays within \(10^{-3}\) of 1, so the learning signal is essentially \(A_s\) alone.</li>
</ul>

<figure id="figure-ppo-gradient-weight">
  <img src="/blog/assets/figures/mental-model-policy-optimizations/ppo-gradient-weight-heatmap.png" alt="PPO gate and learning-signal heatmaps on the advantage–log-ratio plane" />
  <figcaption><strong>Figure 2.</strong> PPO/GRPO. <em>Left:</em> gate \(f'\). <em>Right:</em> learning signal \(f'\rho_t A_t\). Active regions are two quadrant wedges; learning signal grows with \(\rho_t\) in the surviving corners.</figcaption>
</figure>

<figure id="figure-dapo-gradient-weight">
  <img src="/blog/assets/figures/mental-model-policy-optimizations/dapo-gradient-weight-heatmap.png" alt="DAPO gate and learning-signal heatmaps on the advantage–log-ratio plane" />
  <figcaption><strong>Figure 3.</strong> DAPO. Same binary gate as Figure 2, with a wider upper wedge.</figcaption>
</figure>

<figure id="figure-sao-gradient-weight">
  <img src="/blog/assets/figures/mental-model-policy-optimizations/sao-gradient-weight-heatmap.png" alt="SAO gate and learning-signal heatmaps on the advantage–log-ratio plane" />
  <figcaption><strong>Figure 4.</strong> SAO. Top-hat gate. </figcaption>
</figure>

<figure id="figure-sapo-gradient-weight">
  <img src="/blog/assets/figures/mental-model-policy-optimizations/sapo-gradient-weight-heatmap.png" alt="SAPO gate and learning-signal heatmaps on the advantage–log-ratio plane" />
  <figcaption><strong>Figure 5.</strong> SAPO. Smooth gate (dashed lines: \(f'\) iso-levels). <em>Right:</em> over this crop, \(\rho_t\) growth and gate decay nearly cancel.</figcaption>
</figure>

<figure id="figure-gspo-gradient-weight">
  <img src="/blog/assets/figures/mental-model-policy-optimizations/gspo-gradient-weight-heatmap.png" alt="GSPO gradient-weight heatmap on the sequence-advantage / log-sequence-ratio plane" />
  <figcaption><strong>Figure 6.</strong> GSPO on \((A_s,\log\rho_s)\). Same wedge shape as Figure 3, but the axis spans \(\pm10^{-3}\) because GSPO's clip range is three orders of magnitude tighter.</figcaption>
</figure>

<h2 id="takeaways">Takeaways</h2>

<p>Using SAPO’s<sup id="fnref:sapo:3"><a href="#fn:sapo" class="footnote" rel="footnote" role="doc-noteref">6</a></sup> generalized form of objective \(J\) with weight function \(f(\rho_t; A_t)\), learning signals of different RL methods can be shown on the same \((A_t,\log\rho_t)\) plane.</p>

<ul>
  <li>PPO / GRPO: hard token-level gate.</li>
  <li>DAPO: widens the gate for useful low-probability tokens.</li>
  <li>SAO: keeps only ratios inside a trust region when rollout policies become stale.</li>
  <li>SAPO: turns the hard gate into a smooth one.</li>
  <li>GSPO: moves the same idea from tokens to sequences.</li>
</ul>

<p>Seen this way, these methods are providing different answers to the same question: which policy-gradient signals should we trust, and how much?</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:ppo">
      <p>Schulman et al., <a href="https://arxiv.org/abs/1707.06347"><em>Proximal Policy Optimization Algorithms</em></a>, arXiv:1707.06347 (2017). <a href="#fnref:ppo" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:ppo:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a> <a href="#fnref:ppo:2" class="reversefootnote" role="doc-backlink">&#8617;<sup>3</sup></a></p>
    </li>
    <li id="fn:grpo">
      <p>Shao et al., <a href="https://arxiv.org/abs/2402.03300"><em>DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models</em></a>, arXiv:2402.03300 (2024). <a href="#fnref:grpo" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:grpo:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a></p>
    </li>
    <li id="fn:dapo">
      <p>Yu et al., <a href="https://arxiv.org/abs/2503.14476"><em>DAPO: An Open-Source LLM Reinforcement Learning System at Scale</em></a>, arXiv:2503.14476 (2025). <a href="#fnref:dapo" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:dapo:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a></p>
    </li>
    <li id="fn:gspo">
      <p>Zheng et al., <a href="https://arxiv.org/abs/2507.18071"><em>Group Sequence Policy Optimization</em></a>, arXiv:2507.18071 (2025). <a href="#fnref:gspo" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:gspo:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a></p>
    </li>
    <li id="fn:sao">
      <p>Hou et al., <a href="https://arxiv.org/abs/2607.07508"><em>Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning</em></a>, arXiv:2607.07508 (2026). <a href="#fnref:sao" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:sao:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a> <a href="#fnref:sao:2" class="reversefootnote" role="doc-backlink">&#8617;<sup>3</sup></a></p>
    </li>
    <li id="fn:sapo">
      <p>Gao et al., <a href="https://arxiv.org/abs/2511.20347"><em>Soft Adaptive Policy Optimization</em></a>, arXiv:2511.20347 (2025). See also the <a href="https://qwen.ai/blog?id=sapo">Qwen Team blog post on SAPO</a>. <a href="#fnref:sapo" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:sapo:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a> <a href="#fnref:sapo:2" class="reversefootnote" role="doc-backlink">&#8617;<sup>3</sup></a> <a href="#fnref:sapo:3" class="reversefootnote" role="doc-backlink">&#8617;<sup>4</sup></a></p>
    </li>
  </ol>
</div>]]></content><author><name>Zhongwei Wang</name></author><category term="rl" /><category term="ppo" /><category term="grpo" /><category term="gspo" /><category term="dapo" /><category term="sao" /><category term="sapo" /><summary type="html"><![CDATA[Inspired by the Feynman technique, I’ll explain some famous LLM RL losses as simply as I can—starting with PPO1 and GRPO2, then recent variants DAPO3, GSPO4, SAO5, and SAPO6. Writing them out helps me understand them better, and hopefully it helps you too. Schulman et al., Proximal Policy Optimization Algorithms, arXiv:1707.06347 (2017). &#8617; Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, arXiv:2402.03300 (2024). &#8617; Yu et al., DAPO: An Open-Source LLM Reinforcement Learning System at Scale, arXiv:2503.14476 (2025). &#8617; Zheng et al., Group Sequence Policy Optimization, arXiv:2507.18071 (2025). &#8617; Hou et al., Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning, arXiv:2607.07508 (2026). &#8617; Gao et al., Soft Adaptive Policy Optimization, arXiv:2511.20347 (2025). See also the Qwen Team blog post on SAPO. &#8617;]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://wang-zhongwei.github.io/blog/assets/figures/mental-model-policy-optimizations/ppo-clip-schematic.png" /><media:content medium="image" url="https://wang-zhongwei.github.io/blog/assets/figures/mental-model-policy-optimizations/ppo-clip-schematic.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">A Statistical Mechanics View of KL-Regularized RL</title><link href="https://wang-zhongwei.github.io/blog/2026/08/13/a-physicists-dictionary-for-dpo.html" rel="alternate" type="text/html" title="A Statistical Mechanics View of KL-Regularized RL" /><published>2026-08-13T00:00:00+00:00</published><updated>2026-08-13T00:00:00+00:00</updated><id>https://wang-zhongwei.github.io/blog/2026/08/13/a-physicists-dictionary-for-dpo</id><content type="html" xml:base="https://wang-zhongwei.github.io/blog/2026/08/13/a-physicists-dictionary-for-dpo.html"><![CDATA[<p>Preference post-training is often formulated as <strong>KL-regularized reward maximization</strong>: increase expected reward while staying close to a reference policy. Its optimum takes the Gibbs form</p>

\[\pi_\beta^\ast(y|x) = 
\frac{\pi_{\mathrm{ref}}(y|x)e^{r(x,y)/\beta}}{Z(x)},
\tag{1}\]

<p>where \(Z(x)\) is the partition function defined in eq. (13). This same structure underlies DPO’s derivation<sup id="fnref:dpo"><a href="#fn:dpo" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> and KL-regularized RLHF methods such as PPO- and GRPO-style training.<sup id="fnref:ppo"><a href="#fn:ppo" class="footnote" rel="footnote" role="doc-noteref">2</a></sup></p>

<p>Most explanations—including <a href="https://huggingface.co/blog/karina-zadorozhny/guide-to-llm-post-training-algorithms">Karina Zadorozhny’s post-training guide</a> and <a href="https://huggingface.co/blog/ariG23498/rlhf-to-dpo">Ari G’s RLHF-to-DPO walkthrough</a>—essentially stop there. They introduce \(Z(x)\), note its connection to statistical physics, observe that it is intractable but conveniently cancels from the pairwise DPO loss, derive the optimal policy in eq. (1), and move on.</p>

<p><strong>But in statistical mechanics, writing down the partition function is not the end of the derivation. It is the beginning.</strong> Once \(Z\) is defined in eq. (13), free energy, internal energy, entropy, and temperature follow, together with a precise characterization of equilibrium. The KL-regularized RL objective admits an almost term-by-term version of the same structure.</p>

<p>This post works out that dictionary explicitly: reward as negative energy, \(\beta\) as temperature, KL divergence as relative entropy, and the RL objective as a free-energy principle in eq. (4). Then I ask the question the analogy naturally suggests: <strong>if the theory predicts a unique Gibbs optimum \(\pi_\beta^\ast\) in eq. (1), how close does actual training get to it?</strong> I test this with a GRPO-family trainer<sup id="fnref:grpo"><a href="#fn:grpo" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> on GSM8K.</p>

<h2 id="the-kl-regularized-objective">The KL-regularized objective</h2>

<p>Equation (3) in Rafailov et al. (2023)<sup id="fnref:dpo:1"><a href="#fn:dpo" class="footnote" rel="footnote" role="doc-noteref">1</a></sup> defines the KL-regularized RLHF objective that underlies DPO. The same reward–KL tradeoff appears in PPO and GRPO; those methods optimize surrogates or sample-based estimators plus a KL divergence term for regularization with respect to reference policy. In general, for a fixed prompt \(x\), the goal is to optimize</p>

\[\max_{\pi} \; \mathbb{E}_{y\sim\pi(\cdot\mid x)} \left[ r(x,y) \right] - \beta D_{\mathrm{KL}} \left( \pi(\cdot\mid x) \;\|\; \pi_{\mathrm{ref}}(\cdot\mid x) \right).
\tag{2}\]

<p>For readability, we suppress the conditioning on \(x\) and write \(\pi(y)\equiv\pi(y\mid x)\), \(\pi_{\mathrm{ref}}(y)\equiv\pi_{\mathrm{ref}}(y\mid x)\), and \(r(y)\equiv r(x,y)\). The objective becomes</p>

\[J[\pi] = \mathbb{E}_{y\sim\pi}[r(y)] - \beta D_{\mathrm{KL}} \left( \pi \;\|\; \pi_{\mathrm{ref}} \right).
\tag{3}\]

<p>Maximizing \(J[\pi]\) is equivalent to minimizing</p>

\[\boxed{ \mathcal{F}_{\beta}[\pi] = - \mathbb{E}_{y\sim\pi}[r(y)] + \beta D_{\mathrm{KL}} \left( \pi \;\|\; \pi_{\mathrm{ref}} \right). }
\tag{4}\]

<p>Here \(\mathcal{F}_{\beta}\) is a functional of the policy \(\pi\). Given a reference policy \(\pi_{\mathrm{ref}}\) and reward function \(r\), the problem is to find the policy \(\pi\) that minimizes eq. (4).</p>

<h2 id="mapping-to-statistical-mechanics">Mapping to statistical mechanics</h2>

<p>The objective in eq. (4) has the form of a <a href="https://en.wikipedia.org/wiki/Helmholtz_free_energy">Helmholtz free energy</a>. Each response \(y\) is a state of the system, with energy</p>

\[\epsilon(y)=-r(y).
\tag{5}\]

<p>For a policy \(\pi\), define the mean energy and dimensionless relative entropy as</p>

\[U[\pi]=\mathbb{E}_{\pi}[\epsilon]=-\mathbb{E}_{\pi}[r],
\qquad
S_{\mathrm{rel}}[\pi]=-D_{\mathrm{KL}}(\pi\|\pi_{\mathrm{ref}}).
\tag{6}\]

<p>This objective is therefore</p>

\[\mathcal{F}_{\beta}[\pi]
=U[\pi]-\beta S_{\mathrm{rel}}[\pi],
\tag{7}\]

<p>which matches \(F=U-TS\) term by term:</p>

\[\boxed{
\epsilon(y)\longleftrightarrow-r(y),\qquad
k_B T\longleftrightarrow\beta,\qquad
\frac{S}{k_B}\longleftrightarrow-D_{\mathrm{KL}}(\pi\|\pi_{\mathrm{ref}}).
}
\tag{8}\]

<p>Here DPO’s \(\beta\) plays the role of thermal energy \(k_B T\), so its Boltzmann factor \(e^{-\epsilon/\beta}\) matches \(e^{-\beta_{\mathrm{phys}}\epsilon}\) under the physics convention \(\beta_{\mathrm{phys}}=1/(k_B T)\).</p>

<p>But shouldn’t <a href="https://en.wikipedia.org/wiki/Entropy_(statistical_thermodynamics)#Gibbs_entropy_formula">entropy</a> be \(-\sum_{i}\pi_i\log{\pi_i}\)? Why does it involve a reference distribution, \(-\sum_{i}\pi_i\log{\pi_i/\pi_{\mathrm{ref},i}}\)? Ordinary entropy takes exactly this form when states have <a href="https://en.wikipedia.org/wiki/Degenerate_energy_levels">degeneracies</a>. Suppose coarse-grained state \(i\) contains \(g_i\) microstates and has total probability \(\pi_i\). If those microstates are equally likely, each has probability \(p_{i,\alpha}=\pi_i/g_i\), so</p>

\[\frac{S}{k_B}
=-\sum_{i,\alpha}p_{i,\alpha}\log p_{i,\alpha}
=-\sum_i\pi_i\log\frac{\pi_i}{g_i}.\]

<p>Let \(G=\sum_j g_j\) and normalize the degeneracies as \(q_i=g_i/G\). Expanding the logarithm,</p>

\[\frac{S}{k_B}
= -\sum_i \pi_i \log\pi_i + \sum_i \pi_i \log g_i
= -\sum_i \pi_i \log\frac{\pi_i}{q_i} + \log G
= -D_{\mathrm{KL}}(\pi\|q)+\log G.
\tag{9}\]

<p>The constant \(\log G\) does not affect minimization of the free energy in eq. (4). Thus, up to an additive constant, degeneracy turns ordinary entropy into negative KL divergence relative to the normalized degeneracy measure \(q\).</p>

<p>Minimizing \(F=U-TS/k_B=\sum_i\pi_i\epsilon_i+\beta D_{\mathrm{KL}}(\pi\|q)\) subject to \(\sum_i\pi_i=1\) (the same variational problem as in <a href="#b-lagrange-multiplier">Appendix B</a>, with \(\pi_{\mathrm{ref}}\) replaced by \(q\)) gives the Boltzmann distribution</p>

\[\pi_i = \frac{g_i \exp(-\beta_{\mathrm{phys}} \epsilon_i)}{Z},
\qquad
Z = \sum_{i} g_i \exp(-\beta_{\mathrm{phys}} \epsilon_i).
\tag{10}\]

<p>This is the standard form for degenerate systems; see <a href="https://chem.libretexts.org/Bookshelves/Physical_and_Theoretical_Chemistry_Textbook_Maps/Thermodynamics_and_Chemical_Equilibrium_(Ellgen)/21%3A_The_Boltzmann_Distribution_Function/21.01%3A_Finding_the_Boltzmann_Equation">Ellgen, <em>Thermodynamics and Chemical Equilibrium</em>, §21.1</a> and <a href="https://shop.elsevier.com/books/statistical-mechanics/beale/978-0-12-382188-1">Pathria and Beale, <em>Statistical Mechanics</em>, §3.4</a>.</p>

<p>In KL-regularized RL, \(\pi_{\mathrm{ref}}(y)\) is the normalized degeneracy measure \(q(y)\). With \(\epsilon(y)=-r(y)\) from eq. (5) and DPO’s \(\beta\) corresponding to thermal energy \(k_B T\) (so \(\beta_{\mathrm{phys}}=1/\beta\)),</p>

\[e^{-\beta_{\mathrm{phys}}\epsilon}=e^{r/\beta},
\tag{11}\]

<p>which is the Boltzmann weight appearing in eq. (1).</p>

<h2 id="the-gibbs-optimum">The Gibbs optimum</h2>

<p>In statistical mechanics, for fixed temperature, energy levels, and normalized degeneracy measure \(q\), the Gibbs distribution uniquely minimizes the free energy. The same variational principle holds here: once \(\beta\), the reward function, and the reference policy \(\pi_{\mathrm{ref}}\) are fixed, there is a unique optimal policy that minimizes \(\mathcal{F}_{\beta}\) in eq. (4). In a physical system, \(q\) is fixed by degeneracies; in preference optimization, \(\pi_{\mathrm{ref}}\) is a modeling choice that may vary across setups, but within any one setup it serves as the fixed reference.</p>

<p>Following the steps in the <a href="#appendix-two-derivations-of-the-optimal-policy">appendix</a>, the free energy decomposes as</p>

\[\boxed{ \mathcal{F}_{\beta}[\pi] = -\beta\log Z_{\beta} + \beta D_{\mathrm{KL}} \left( \pi \;\|\; \pi_{\beta}^{\ast} \right). }
\tag{12}\]

<p>where the partition function is</p>

\[\boxed{ Z_{\beta} = \sum_y \pi_{\mathrm{ref}}(y) e^{r(y)/\beta} }
\tag{13}\]

<p>and the corresponding Gibbs policy is eq. (1), with \(Z(x)\) replaced by \(Z_{\beta}\) when conditioning on \(x\) is suppressed.</p>

<p>Since \(D_{\mathrm{KL}}(\pi\|\pi_{\beta}^{\ast})\ge 0\), with equality if and only if \(\pi=\pi_{\beta}^{\ast}\), the unique minimizer is \(\pi_{\beta}^{\ast}\) and the equilibrium free energy is \(\mathcal{F}_{\beta}^{\ast}=-\beta\log Z_{\beta}\). This is exactly the free energy of a <a href="https://en.wikipedia.org/wiki/Helmholtz_free_energy#Relation_to_the_canonical_partition_function">canonical ensemble</a>. More generally, once the <a href="https://en.wikipedia.org/wiki/Partition_function_(statistical_mechanics)#Relation_to_thermodynamic_variables">partition function</a> in eq. (13) is known, equilibrium quantities such as internal energy, entropy, and heat capacity follow from it and its derivatives with respect to \(\beta\)—the same logic summarized in the dictionary below and in eq. (8).</p>

<p>There are two standard ways to derive \(\pi_{\beta}^{\ast}\):</p>

<ol>
  <li><strong>Partition-function / KL decomposition.</strong> Rewrite eq. (4) so that an unnormalized Boltzmann weight appears, then normalize it by \(Z_{\beta}\) from eq. (13).</li>
  <li><strong>Lagrange multiplier.</strong> Enforce the normalization constraint \(\sum_y\pi(y)=1\) while setting the functional derivative of \(\mathcal{F}_{\beta}\) to zero; the stationary point is again eq. (1).</li>
</ol>

<p>Full algebra for both routes is in the <a href="#appendix-two-derivations-of-the-optimal-policy">appendix</a>.</p>

<h2 id="the-dictionary">The dictionary</h2>

<p>Statistical mechanics on the left, KL-regularized RL on the right:</p>

<table>
  <thead>
    <tr>
      <th>Statistical mechanics</th>
      <th>KL-regularized RL</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Energy of state \(\epsilon\)</td>
      <td>\(-r\) (negative reward)</td>
    </tr>
    <tr>
      <td>Normalized degeneracy \(q\)</td>
      <td>\(\pi_{\mathrm{ref}}\)</td>
    </tr>
    <tr>
      <td>Internal energy \(U\)</td>
      <td>\(-\mathbb{E}_{\pi}[r]\)</td>
    </tr>
    <tr>
      <td>Entropy \(S\)</td>
      <td>\(-D_{\mathrm{KL}}(\pi \Vert \pi_{\mathrm{ref}})\)</td>
    </tr>
    <tr>
      <td>Temperature \(k_B T\)</td>
      <td>DPO’s \(\beta\)</td>
    </tr>
    <tr>
      <td>Partition function \(Z\)</td>
      <td>\(\sum_y \pi_{\mathrm{ref}}(y)\,e^{r(y)/\beta}\)</td>
    </tr>
    <tr>
      <td>Minimum free energy \(F^{\ast}\)</td>
      <td>\(-\beta \log Z_\beta\)</td>
    </tr>
    <tr>
      <td>Free-energy gap</td>
      <td>\(\beta\, D_{\mathrm{KL}}(\pi \Vert \pi_\beta^{\ast})\)</td>
    </tr>
    <tr>
      <td>Heat capacity \(C_V/k_B\)</td>
      <td>\(\displaystyle \frac{\operatorname{Var}_{\pi_{\beta}^{\ast}}[r]}{\beta^2} = -\frac{\partial \mathbb{E}_{\pi_{\beta}^{\ast}}[r]}{\partial \beta}\)</td>
    </tr>
  </tbody>
</table>

<p>The heat-capacity entry follows from differentiating the partition function in eq. (13). Write \(t=1/\beta\) so that \(Z_{\beta}=\sum_y\pi_{\mathrm{ref}}(y)e^{tr(y)}\). Standard identities for the Gibbs policy in eq. (1) give \(\mathbb{E}_{\pi_{\beta}^{\ast}}[r]=\partial_t\log Z_{\beta}\) and \(\operatorname{Var}_{\pi_{\beta}^{\ast}}[r]=\partial_t^2\log Z_{\beta}\). Since \(t=1/\beta\) implies \(\mathrm{d}t/\mathrm{d}\beta=-1/\beta^2\),</p>

\[-\frac{\partial \mathbb{E}_{\pi_{\beta}^{\ast}}[r]}{\partial \beta}
= -\frac{\mathrm{d}t}{\mathrm{d}\beta}\,\frac{\partial \mathbb{E}_{\pi_{\beta}^{\ast}}[r]}{\partial t}
= \frac{\operatorname{Var}_{\pi_{\beta}^{\ast}}[r]}{\beta^2}.\]

<p>With \(\epsilon=-r\), this is the usual canonical-ensemble relation \(C_V/k_B=\beta_{\mathrm{phys}}^2\operatorname{Var}(\epsilon)\).</p>

<h2 id="experiments-on-gsm8k">Experiments on GSM8K</h2>

<p>Everything above is exact at the level of probability distributions. To see what that ideal predicts for a real model, I ran <code class="language-plaintext highlighter-rouge">Qwen2.5-0.5B-Instruct</code> on GSM8K with a binary reward: \(r(y)=1\) when the final answer is correct and \(0\) otherwise.</p>

<h3 id="the-gibbs-floor">The Gibbs floor</h3>

<p>For each of 64 GSM8K test prompts \(x_i\), I sampled 24 candidate responses \(y_{ij}\) from the base model, giving 1,536 completions in total. Generation used temperature 0.8, top-p 0.95, and a 512-token limit. Each candidate has a base-model log-probability and a binary correctness reward \(r(x_i,y_{ij})\in\{0,1\}\). Normalizing the base-model probabilities over the 24 candidates defines \(\pi_{\mathrm{ref}}(y_{ij}\mid x_i)\) on this finite candidate set.</p>

<p>Treating these 24 candidates as the state space turns every expectation into a finite sum. For each prompt, eqs. (1) and (13) become</p>

\[Z_{\beta}(x_i)
=\sum_{j=1}^{24}\pi_{\mathrm{ref}}(y_{ij}\mid x_i)
e^{r(x_i,y_{ij})/\beta},
\qquad
\pi_{\beta}^{\ast}(y_{ij}\mid x_i)
=\frac{\pi_{\mathrm{ref}}(y_{ij}\mid x_i)e^{r(x_i,y_{ij})/\beta}}
{Z_{\beta}(x_i)}.
\tag{14}\]

<p>Once the candidate set has been sampled, the Gibbs policy and its expected reward, KL divergence, and free energy from eq. (4) can all be evaluated exactly on that set. Sweeping \(\beta\) traces the optimal reward–KL frontier in <a href="#figure-reward-kl-frontier">Figure 1</a> and the corresponding free-energy landscape in <a href="#figure-free-energy-landscape">Figure 2</a>. Figure 1 also compares two ways of defining the reference weights: full-sequence probability and length-normalized probability. The remaining experiments use full-sequence probability.</p>

<figure id="figure-reward-kl-frontier">
  <img src="/blog/assets/figures/a-physicists-dictionary-for-dpo/reward-kl-frontier.png" alt="Reward–KL frontier across beta values" />
  <figcaption><strong>Figure 1.</strong> Reward–KL frontier across \(\beta\), comparing sequence probability with length-normalized probability.</figcaption>
</figure>

<figure id="figure-free-energy-landscape">
  <img src="/blog/assets/figures/a-physicists-dictionary-for-dpo/free-energy-landscape.png" alt="Free-energy landscape across beta values" />
  <figcaption><strong>Figure 2.</strong> Free-energy landscape across \(\beta\); color indicates the KL-based gap to the Gibbs optimum.</figcaption>
</figure>

<h3 id="grpo-family-training-vs-the-floor">GRPO-family training vs. the floor</h3>

<p>At the reference policy the KL term is zero, so \(\mathcal{F}_{\beta}[\pi_{\mathrm{ref}}]=-\mathbb{E}_{\pi_{\mathrm{ref}}}[r]\), shown as the black dotted baseline in <a href="#figure-free-energy-landscape">Figure 2</a>. If training reaches the distribution-level optimum, its free energy should fall from this baseline toward the red Gibbs minimum at \(\mathcal{F}_{\beta}^{\ast}=-\beta\log Z_{\beta}\) from eq. (12).</p>

<p>To test this, I ran a GRPO-family trainer<sup id="fnref:grpo:1"><a href="#fn:grpo" class="footnote" rel="footnote" role="doc-noteref">3</a></sup> separately for \(\beta\in\{0.01,0.02,0.05,0.1,0.2\}\), using group size \(G=32\), 6,500 GSM8K training prompts, for about one epoch (262,144 rollouts). I saved checkpoints every 32,768 rollouts and evaluated each checkpoint on the same 500 held-out prompts. The reference model starts at an expected reward of 0.348.</p>

<figure id="figure-training-free-energy-decomposition">
  <img src="/blog/assets/figures/a-physicists-dictionary-for-dpo/training-free-energy-decomposition.png" alt="Training-time free-energy decomposition across KL penalty values" />
  <figcaption><strong>Figure 3.</strong> Training-time free-energy decomposition across KL penalties \(\beta\): reward, KL, and total free-energy terms on held-out validation.</figcaption>
</figure>

<h3 id="why-training-falls-short">Why training falls short</h3>

<p>I expected each policy to move toward its \(\pi_{\beta}^{\ast}\) from eq. (1) and each free-energy curve to approach its dashed target from eq. (12). <a href="#figure-training-free-energy-decomposition">Figure 3</a> shows that this was too optimistic: after one epoch, none of the runs reached its Gibbs minimum. Reward improved, but not enough to offset the accompanying \(\beta\)-weighted KL cost, so the measured free energy stayed above the target.</p>

<p>Likely explanations:</p>

<ol>
  <li><strong>Limited expressivity.</strong> Optimization is restricted to the parameter space of a 0.5B model, which may not contain the true Gibbs policy.</li>
  <li><strong>Non-convex loss.</strong> Even within that parameter space, the loss is non-convex, so gradient descent is not guaranteed to find the global minimum.</li>
</ol>

<h2 id="takeaways">Takeaways</h2>

<ul>
  <li>KL-regularized RL is free-energy minimization in statistical physics: eq. (4) is the RL free energy and eq. (8) is the dictionary.
    <ul>
      <li>Given \(\beta\), reference policy \(\pi_{\mathrm{ref}}\), and reward function \(r\), there is a unique policy \(\pi_{\beta}^{\ast}\) in eq. (1) that minimizes the free energy.</li>
      <li>From \(Z_{\beta}\) in eq. (13) we can read off internal energy, entropy, temperature, free energy, and heat capacity.</li>
    </ul>
  </li>
  <li>Training may not reach that distribution-level optimum.</li>
</ul>

<p>Code and full experiments: <a href="https://github.com/wang-zhongwei/stat-mech-dpo">github.com/wang-zhongwei/stat-mech-dpo</a></p>

<h2 id="appendix-two-derivations-of-the-optimal-policy">Appendix: two derivations of the optimal policy</h2>

<p>Both routes start from eq. (4), equivalently written as</p>

\[\boxed{\mathcal{F}_{\beta}[\pi] = -\sum_y \pi(y)r(y) + \beta\sum_y\pi(y)\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)}}
\tag{A.0}\]

<p>and the normalization constraint \(\sum_y\pi(y)=1\). They recover the partition function in eq. (13), the optimal policy in eq. (1), and the free-energy minimum in eq. (12).</p>

<h3 id="a-partition-function--kl-decomposition">A. Partition-function / KL decomposition</h3>

<p>Writing the expectation and KL divergence explicitly and dividing by \(\beta\) gives</p>

\[\frac{\mathcal{F}_{\beta}[\pi]}{\beta}
= \sum_y \pi(y)\left[\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)}-\frac{r(y)}{\beta}\right]
= \sum_y \pi(y)\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)e^{r(y)/\beta}}.
\tag{A.1}\]

<p>The Boltzmann weight \(\pi_{\mathrm{ref}}(y)e^{r(y)/\beta}\) from eq. (11) appears directly from the objective—not as an ansatz. Normalizing it by the partition function in eq. (13) gives the optimal policy in eq. (1).</p>

<p>Substituting \(\pi_{\mathrm{ref}}(y)e^{r(y)/\beta}=Z_{\beta}\pi_{\beta}^{\ast}(y)\) then yields</p>

\[\frac{\mathcal{F}_{\beta}[\pi]}{\beta}
= \sum_y \pi(y)\log\frac{\pi(y)}{Z_{\beta}\pi_{\beta}^{\ast}(y)}
= D_{\mathrm{KL}}\left(\pi\;\|\;\pi_{\beta}^{\ast}\right)
- \log Z_{\beta}\sum_y\pi(y).
\tag{A.2}\]

<p>Because \(\pi\) is a normalized probability distribution, \(\sum_y\pi(y)=1\). Multiplying by \(\beta\) therefore gives eq. (12):</p>

\[\mathcal{F}_{\beta}[\pi] = -\beta\log Z_{\beta} + \beta D_{\mathrm{KL}}\left(\pi\;\|\;\pi_{\beta}^{\ast}\right).
\tag{A.3}\]

<p>Since \(D_{\mathrm{KL}}(\pi\|\pi_{\beta}^{\ast})\ge 0\), with equality if and only if \(\pi=\pi_{\beta}^{\ast}\), the unique minimizer is eq. (1) and \(\mathcal{F}_{\beta}^{\ast}=-\beta\log Z_{\beta}\).</p>

<p>With \(\epsilon=-r\) from eq. (5), the same algebra produces the statistical-mechanics form</p>

\[\pi_{\beta}^{\ast}(y)=\frac{\pi_{\mathrm{ref}}(y)e^{-\epsilon(y)/\beta}}{Z_{\beta}},\qquad Z_{\beta}=\sum_y\pi_{\mathrm{ref}}(y)e^{-\epsilon(y)/\beta}.
\tag{A.4}\]

<p>For a uniform \(\pi_{\mathrm{ref}}\), this reduces to the canonical Gibbs distribution \(e^{-\epsilon/\beta}/Z_{\beta}\).</p>

<h3 id="b-lagrange-multiplier">B. Lagrange multiplier</h3>

<p>To minimize eq. (4) subject to \(\sum_y\pi(y)=1\), introduce a Lagrange multiplier \(\lambda\):</p>

\[\mathcal{L}[\pi,\lambda]
= -\sum_y\pi(y)r(y)
+\beta\sum_y\pi(y)\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)}
+\lambda\left(\sum_y\pi(y)-1\right).
\tag{B.1}\]

<p>Stationarity with respect to each \(\pi(y)\) requires</p>

\[\frac{\partial\mathcal{L}}{\partial\pi(y)}
= -r(y)
+\beta\left(\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)}+1\right)
+\lambda
=0.
\tag{B.2}\]

<p>Solving for \(\pi(y)\) gives</p>

\[\pi(y)=C\,\pi_{\mathrm{ref}}(y)e^{r(y)/\beta},
\tag{B.3}\]

<p>where \(C=e^{-1-\lambda/\beta}\) is independent of \(y\). Enforcing normalization yields \(C^{-1}=Z_{\beta}\) from eq. (13), and therefore the same optimal policy as eq. (1):</p>

\[\pi_{\beta}^{\ast}(y)=\frac{\pi_{\mathrm{ref}}(y)e^{r(y)/\beta}}{Z_{\beta}}.
\tag{B.4}\]

<p>Substituting this stationary point back into eq. (4) recovers the equilibrium free energy \(\mathcal{F}_{\beta}^{\ast}=-\beta\log Z_{\beta}\) from eq. (12).</p>

<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:dpo">
      <p>Rafailov et al., <a href="https://proceedings.neurips.cc/paper_files/paper/2023/file/a85b405ed65c6477a4fe8302b5e06ce7-Paper-Conference.pdf"><em>Direct Preference Optimization: Your Language Model Is Secretly a Reward Model</em></a>, NeurIPS 2023. <a href="#fnref:dpo" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:dpo:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a></p>
    </li>
    <li id="fn:ppo">
      <p>Schulman et al., <a href="https://arxiv.org/abs/1707.06347"><em>Proximal Policy Optimization Algorithms</em></a>, arXiv:1707.06347 (2017). PPO does not optimize the KL-regularized objective in eq. (2) directly: it maximizes a clipped surrogate in the ratio \(\rho=\pi_\theta/\pi_{\theta_{\mathrm{old}}}\). But at the start of each update, where \(\pi_\theta=\pi_{\theta_{\mathrm{old}}}\), the clip is inactive and \(\nabla_\theta\,\rho A=\nabla_\theta \log\pi_\theta\,A\)—the vanilla policy gradient. With the KL-to-reference penalty folded into the reward, PPO therefore locally ascends the same KL-regularized objective; clipping only limits the step size. <a href="#fnref:ppo" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:grpo">
      <p>Shao et al., <a href="https://arxiv.org/abs/2402.03300"><em>DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models</em></a>, arXiv:2402.03300 (2024). <a href="#fnref:grpo" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:grpo:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a></p>
    </li>
  </ol>
</div>]]></content><author><name>Zhongwei Wang</name></author><category term="statistical-mechanics" /><category term="rlhf" /><category term="dpo" /><summary type="html"><![CDATA[Preference post-training is often formulated as KL-regularized reward maximization: increase expected reward while staying close to a reference policy. Its optimum takes the Gibbs form]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://wang-zhongwei.github.io/blog/assets/figures/a-physicists-dictionary-for-dpo/free-energy-landscape.png" /><media:content medium="image" url="https://wang-zhongwei.github.io/blog/assets/figures/a-physicists-dictionary-for-dpo/free-energy-landscape.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>