Preference post-training is often formulated as KL-regularized reward maximization: increase expected reward while staying close to a reference policy. Its optimum takes the Gibbs form

πβ∗(y∣x)=πref(y∣x)er(x,y)/βZ(x),(1)\pi_\beta^\ast(y|x) = \frac{\pi_{\mathrm{ref}}(y|x)e^{r(x,y)/\beta}}{Z(x)}, \tag{1}

where Z(x)Z(x) is the partition function defined in eq. (13). This same structure underlies DPO’s derivation1 and KL-regularized RLHF methods such as PPO- and GRPO-style training.2

Most explanations—including Karina Zadorozhny’s post-training guide and Ari G’s RLHF-to-DPO walkthrough—essentially stop there. They introduce Z(x)Z(x), note its connection to statistical physics, observe that it is intractable but conveniently cancels from the pairwise DPO loss, derive the optimal policy in eq. (1), and move on.

But in statistical mechanics, writing down the partition function is not the end of the derivation. It is the beginning. Once ZZ is defined in eq. (13), free energy, internal energy, entropy, and temperature follow, together with a precise characterization of equilibrium. The KL-regularized RL objective admits an almost term-by-term version of the same structure.

This post works out that dictionary explicitly: reward as negative energy, β\beta as temperature, KL divergence as relative entropy, and the RL objective as a free-energy principle in eq. (4). Then I ask the question the analogy naturally suggests: if the theory predicts a unique Gibbs optimum πβ∗\pi_\beta^\ast in eq. (1), how close does actual training get to it? I test this with a GRPO-family trainer3 on GSM8K.

The KL-regularized objective

Equation (3) in Rafailov et al. (2023)1 defines the KL-regularized RLHF objective that underlies DPO. The same reward–KL tradeoff appears in PPO and GRPO; those methods optimize surrogates or sample-based estimators plus a KL divergence term for regularization with respect to reference policy. In general, for a fixed prompt xx, the goal is to optimize

max⁡π  Ey∼π(⋅∣x)[r(x,y)]−βDKL(π(⋅∣x)  ∥  πref(⋅∣x)).(2)\max_{\pi} \; \mathbb{E}_{y\sim\pi(\cdot\mid x)} \left[ r(x,y) \right] - \beta D_{\mathrm{KL}} \left( \pi(\cdot\mid x) \;\|\; \pi_{\mathrm{ref}}(\cdot\mid x) \right). \tag{2}

For readability, we suppress the conditioning on xx and write π(y)≡π(y∣x)\pi(y)\equiv\pi(y\mid x), πref(y)≡πref(y∣x)\pi_{\mathrm{ref}}(y)\equiv\pi_{\mathrm{ref}}(y\mid x), and r(y)≡r(x,y)r(y)\equiv r(x,y). The objective becomes

J[π]=Ey∼π[r(y)]−βDKL(π  ∥  πref).(3)J[\pi] = \mathbb{E}_{y\sim\pi}[r(y)] - \beta D_{\mathrm{KL}} \left( \pi \;\|\; \pi_{\mathrm{ref}} \right). \tag{3}

Maximizing J[π]J[\pi] is equivalent to minimizing

Fβ[π]=−Ey∼π[r(y)]+βDKL(π  ∥  πref).(4)\boxed{ \mathcal{F}_{\beta}[\pi] = - \mathbb{E}_{y\sim\pi}[r(y)] + \beta D_{\mathrm{KL}} \left( \pi \;\|\; \pi_{\mathrm{ref}} \right). } \tag{4}

Here Fβ\mathcal{F}_{\beta} is a functional of the policy π\pi. Given a reference policy πref\pi_{\mathrm{ref}} and reward function rr, the problem is to find the policy π\pi that minimizes eq. (4).

Mapping to statistical mechanics

The objective in eq. (4) has the form of a Helmholtz free energy. Each response yy is a state of the system, with energy

ϵ(y)=−r(y).(5)\epsilon(y)=-r(y). \tag{5}

For a policy π\pi, define the mean energy and dimensionless relative entropy as

U[π]=Eπ[ϵ]=−Eπ[r],Srel[π]=−DKL(π∥πref).(6)U[\pi]=\mathbb{E}_{\pi}[\epsilon]=-\mathbb{E}_{\pi}[r], \qquad S_{\mathrm{rel}}[\pi]=-D_{\mathrm{KL}}(\pi\|\pi_{\mathrm{ref}}). \tag{6}

This objective is therefore

Fβ[π]=U[π]−βSrel[π],(7)\mathcal{F}_{\beta}[\pi] =U[\pi]-\beta S_{\mathrm{rel}}[\pi], \tag{7}

which matches F=U−TSF=U-TS term by term:

ϵ(y)⟷−r(y),kBT⟷β,SkB⟷−DKL(π∥πref).(8)\boxed{ \epsilon(y)\longleftrightarrow-r(y),\qquad k_B T\longleftrightarrow\beta,\qquad \frac{S}{k_B}\longleftrightarrow-D_{\mathrm{KL}}(\pi\|\pi_{\mathrm{ref}}). } \tag{8}

Here DPO’s β\beta plays the role of thermal energy kBTk_B T, so its Boltzmann factor e−ϵ/βe^{-\epsilon/\beta} matches e−βphysϵe^{-\beta_{\mathrm{phys}}\epsilon} under the physics convention βphys=1/(kBT)\beta_{\mathrm{phys}}=1/(k_B T).

But shouldn’t entropy be −∑iπilog⁡πi-\sum_{i}\pi_i\log{\pi_i}? Why does it involve a reference distribution, −∑iπilog⁡πi/πref,i-\sum_{i}\pi_i\log{\pi_i/\pi_{\mathrm{ref},i}}? Ordinary entropy takes exactly this form when states have degeneracies. Suppose coarse-grained state ii contains gig_i microstates and has total probability πi\pi_i. If those microstates are equally likely, each has probability pi,α=πi/gip_{i,\alpha}=\pi_i/g_i, so

SkB=−∑i,αpi,αlog⁡pi,α=−∑iπilog⁡πigi.\frac{S}{k_B} =-\sum_{i,\alpha}p_{i,\alpha}\log p_{i,\alpha} =-\sum_i\pi_i\log\frac{\pi_i}{g_i}.

Let G=∑jgjG=\sum_j g_j and normalize the degeneracies as qi=gi/Gq_i=g_i/G. Expanding the logarithm,

SkB=−∑iπilog⁡πi+∑iπilog⁡gi=−∑iπilog⁡πiqi+log⁡G=−DKL(π∥q)+log⁡G.(9)\frac{S}{k_B} = -\sum_i \pi_i \log\pi_i + \sum_i \pi_i \log g_i = -\sum_i \pi_i \log\frac{\pi_i}{q_i} + \log G = -D_{\mathrm{KL}}(\pi\|q)+\log G. \tag{9}

The constant log⁡G\log G does not affect minimization of the free energy in eq. (4). Thus, up to an additive constant, degeneracy turns ordinary entropy into negative KL divergence relative to the normalized degeneracy measure qq.

Minimizing F=U−TS/kB=∑iπiϵi+βDKL(π∥q)F=U-TS/k_B=\sum_i\pi_i\epsilon_i+\beta D_{\mathrm{KL}}(\pi\|q) subject to ∑iπi=1\sum_i\pi_i=1 (the same variational problem as in Appendix B, with πref\pi_{\mathrm{ref}} replaced by qq) gives the Boltzmann distribution

πi=giexp⁡(−βphysϵi)Z,Z=∑igiexp⁡(−βphysϵi).(10)\pi_i = \frac{g_i \exp(-\beta_{\mathrm{phys}} \epsilon_i)}{Z}, \qquad Z = \sum_{i} g_i \exp(-\beta_{\mathrm{phys}} \epsilon_i). \tag{10}

This is the standard form for degenerate systems; see Ellgen, Thermodynamics and Chemical Equilibrium, §21.1 and Pathria and Beale, Statistical Mechanics, §3.4.

In KL-regularized RL, πref(y)\pi_{\mathrm{ref}}(y) is the normalized degeneracy measure q(y)q(y). With ϵ(y)=−r(y)\epsilon(y)=-r(y) from eq. (5) and DPO’s β\beta corresponding to thermal energy kBTk_B T (so βphys=1/β\beta_{\mathrm{phys}}=1/\beta),

e−βphysϵ=er/β,(11)e^{-\beta_{\mathrm{phys}}\epsilon}=e^{r/\beta}, \tag{11}

which is the Boltzmann weight appearing in eq. (1).

The Gibbs optimum

In statistical mechanics, for fixed temperature, energy levels, and normalized degeneracy measure qq, the Gibbs distribution uniquely minimizes the free energy. The same variational principle holds here: once β\beta, the reward function, and the reference policy πref\pi_{\mathrm{ref}} are fixed, there is a unique optimal policy that minimizes Fβ\mathcal{F}_{\beta} in eq. (4). In a physical system, qq is fixed by degeneracies; in preference optimization, πref\pi_{\mathrm{ref}} is a modeling choice that may vary across setups, but within any one setup it serves as the fixed reference.

Following the steps in the appendix, the free energy decomposes as

Fβ[π]=−βlog⁡Zβ+βDKL(π  ∥  πβ∗).(12)\boxed{ \mathcal{F}_{\beta}[\pi] = -\beta\log Z_{\beta} + \beta D_{\mathrm{KL}} \left( \pi \;\|\; \pi_{\beta}^{\ast} \right). } \tag{12}

where the partition function is

Zβ=∑yπref(y)er(y)/β(13)\boxed{ Z_{\beta} = \sum_y \pi_{\mathrm{ref}}(y) e^{r(y)/\beta} } \tag{13}

and the corresponding Gibbs policy is eq. (1), with Z(x)Z(x) replaced by ZβZ_{\beta} when conditioning on xx is suppressed.

Since DKL(π∥πβ∗)≥0D_{\mathrm{KL}}(\pi\|\pi_{\beta}^{\ast})\ge 0, with equality if and only if π=πβ∗\pi=\pi_{\beta}^{\ast}, the unique minimizer is πβ∗\pi_{\beta}^{\ast} and the equilibrium free energy is Fβ∗=−βlog⁡Zβ\mathcal{F}_{\beta}^{\ast}=-\beta\log Z_{\beta}. This is exactly the free energy of a canonical ensemble. More generally, once the partition function in eq. (13) is known, equilibrium quantities such as internal energy, entropy, and heat capacity follow from it and its derivatives with respect to β\beta—the same logic summarized in the dictionary below and in eq. (8).

There are two standard ways to derive πβ∗\pi_{\beta}^{\ast}:

  1. Partition-function / KL decomposition. Rewrite eq. (4) so that an unnormalized Boltzmann weight appears, then normalize it by ZβZ_{\beta} from eq. (13).
  2. Lagrange multiplier. Enforce the normalization constraint ∑yπ(y)=1\sum_y\pi(y)=1 while setting the functional derivative of Fβ\mathcal{F}_{\beta} to zero; the stationary point is again eq. (1).

Full algebra for both routes is in the appendix.

The dictionary

Statistical mechanics on the left, KL-regularized RL on the right:

Statistical mechanics KL-regularized RL
Energy of state ϵ\epsilon −r-r (negative reward)
Normalized degeneracy qq πref\pi_{\mathrm{ref}}
Internal energy UU −Eπ[r]-\mathbb{E}_{\pi}[r]
Entropy SS −DKL(π∥πref)-D_{\mathrm{KL}}(\pi \Vert \pi_{\mathrm{ref}})
Temperature kBTk_B T DPO’s β\beta
Partition function ZZ ∑yπref(y) er(y)/β\sum_y \pi_{\mathrm{ref}}(y)\,e^{r(y)/\beta}
Minimum free energy F∗F^{\ast} −βlog⁡Zβ-\beta \log Z_\beta
Free-energy gap β DKL(π∥πβ∗)\beta\, D_{\mathrm{KL}}(\pi \Vert \pi_\beta^{\ast})
Heat capacity CV/kBC_V/k_B Var⁡πβ∗[r]β2=−∂Eπβ∗[r]∂β\displaystyle \frac{\operatorname{Var}_{\pi_{\beta}^{\ast}}[r]}{\beta^2} = -\frac{\partial \mathbb{E}_{\pi_{\beta}^{\ast}}[r]}{\partial \beta}

The heat-capacity entry follows from differentiating the partition function in eq. (13). Write t=1/βt=1/\beta so that Zβ=∑yπref(y)etr(y)Z_{\beta}=\sum_y\pi_{\mathrm{ref}}(y)e^{tr(y)}. Standard identities for the Gibbs policy in eq. (1) give Eπβ∗[r]=∂tlog⁡Zβ\mathbb{E}_{\pi_{\beta}^{\ast}}[r]=\partial_t\log Z_{\beta} and Var⁡πβ∗[r]=∂t2log⁡Zβ\operatorname{Var}_{\pi_{\beta}^{\ast}}[r]=\partial_t^2\log Z_{\beta}. Since t=1/βt=1/\beta implies dt/dβ=−1/β2\mathrm{d}t/\mathrm{d}\beta=-1/\beta^2,

−∂Eπβ∗[r]∂β=−dtdβ ∂Eπβ∗[r]∂t=Var⁡πβ∗[r]β2.-\frac{\partial \mathbb{E}_{\pi_{\beta}^{\ast}}[r]}{\partial \beta} = -\frac{\mathrm{d}t}{\mathrm{d}\beta}\,\frac{\partial \mathbb{E}_{\pi_{\beta}^{\ast}}[r]}{\partial t} = \frac{\operatorname{Var}_{\pi_{\beta}^{\ast}}[r]}{\beta^2}.

With ϵ=−r\epsilon=-r, this is the usual canonical-ensemble relation CV/kB=βphys2Var⁡(ϵ)C_V/k_B=\beta_{\mathrm{phys}}^2\operatorname{Var}(\epsilon).

Experiments on GSM8K

Everything above is exact at the level of probability distributions. To see what that ideal predicts for a real model, I ran Qwen2.5-0.5B-Instruct on GSM8K with a binary reward: r(y)=1r(y)=1 when the final answer is correct and 00 otherwise.

The Gibbs floor

For each of 64 GSM8K test prompts xix_i, I sampled 24 candidate responses yijy_{ij} from the base model, giving 1,536 completions in total. Generation used temperature 0.8, top-p 0.95, and a 512-token limit. Each candidate has a base-model log-probability and a binary correctness reward r(xi,yij)∈{0,1}r(x_i,y_{ij})\in\{0,1\}. Normalizing the base-model probabilities over the 24 candidates defines πref(yij∣xi)\pi_{\mathrm{ref}}(y_{ij}\mid x_i) on this finite candidate set.

Treating these 24 candidates as the state space turns every expectation into a finite sum. For each prompt, eqs. (1) and (13) become

Zβ(xi)=∑j=124πref(yij∣xi)er(xi,yij)/β,πβ∗(yij∣xi)=πref(yij∣xi)er(xi,yij)/βZβ(xi).(14)Z_{\beta}(x_i) =\sum_{j=1}^{24}\pi_{\mathrm{ref}}(y_{ij}\mid x_i) e^{r(x_i,y_{ij})/\beta}, \qquad \pi_{\beta}^{\ast}(y_{ij}\mid x_i) =\frac{\pi_{\mathrm{ref}}(y_{ij}\mid x_i)e^{r(x_i,y_{ij})/\beta}} {Z_{\beta}(x_i)}. \tag{14}

Once the candidate set has been sampled, the Gibbs policy and its expected reward, KL divergence, and free energy from eq. (4) can all be evaluated exactly on that set. Sweeping β\beta traces the optimal reward–KL frontier in Figure 1 and the corresponding free-energy landscape in Figure 2. Figure 1 also compares two ways of defining the reference weights: full-sequence probability and length-normalized probability. The remaining experiments use full-sequence probability.

Reward–KL frontier across beta values
Figure 1. Reward–KL frontier across β\beta, comparing sequence probability with length-normalized probability.
Free-energy landscape across beta values
Figure 2. Free-energy landscape across β\beta; color indicates the KL-based gap to the Gibbs optimum.

GRPO-family training vs. the floor

At the reference policy the KL term is zero, so Fβ[πref]=−Eπref[r]\mathcal{F}_{\beta}[\pi_{\mathrm{ref}}]=-\mathbb{E}_{\pi_{\mathrm{ref}}}[r], shown as the black dotted baseline in Figure 2. If training reaches the distribution-level optimum, its free energy should fall from this baseline toward the red Gibbs minimum at Fβ∗=−βlog⁡Zβ\mathcal{F}_{\beta}^{\ast}=-\beta\log Z_{\beta} from eq. (12).

To test this, I ran a GRPO-family trainer3 separately for β∈{0.01,0.02,0.05,0.1,0.2}\beta\in\{0.01,0.02,0.05,0.1,0.2\}, using group size G=32G=32, 6,500 GSM8K training prompts, for about one epoch (262,144 rollouts). I saved checkpoints every 32,768 rollouts and evaluated each checkpoint on the same 500 held-out prompts. The reference model starts at an expected reward of 0.348.

Training-time free-energy decomposition across KL penalty values
Figure 3. Training-time free-energy decomposition across KL penalties β\beta: reward, KL, and total free-energy terms on held-out validation.

Why training falls short

I expected each policy to move toward its πβ∗\pi_{\beta}^{\ast} from eq. (1) and each free-energy curve to approach its dashed target from eq. (12). Figure 3 shows that this was too optimistic: after one epoch, none of the runs reached its Gibbs minimum. Reward improved, but not enough to offset the accompanying β\beta-weighted KL cost, so the measured free energy stayed above the target.

Likely explanations:

  1. Limited expressivity. Optimization is restricted to the parameter space of a 0.5B model, which may not contain the true Gibbs policy.
  2. Non-convex loss. Even within that parameter space, the loss is non-convex, so gradient descent is not guaranteed to find the global minimum.

Takeaways

  • KL-regularized RL is free-energy minimization in statistical physics: eq. (4) is the RL free energy and eq. (8) is the dictionary.
    • Given β\beta, reference policy πref\pi_{\mathrm{ref}}, and reward function rr, there is a unique policy πβ∗\pi_{\beta}^{\ast} in eq. (1) that minimizes the free energy.
    • From ZβZ_{\beta} in eq. (13) we can read off internal energy, entropy, temperature, free energy, and heat capacity.
  • Training may not reach that distribution-level optimum.

Code and full experiments: github.com/wang-zhongwei/stat-mech-dpo

Appendix: two derivations of the optimal policy

Both routes start from eq. (4), equivalently written as

Fβ[π]=−∑yπ(y)r(y)+β∑yπ(y)log⁡π(y)πref(y)(A.0)\boxed{\mathcal{F}_{\beta}[\pi] = -\sum_y \pi(y)r(y) + \beta\sum_y\pi(y)\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)}} \tag{A.0}

and the normalization constraint ∑yπ(y)=1\sum_y\pi(y)=1. They recover the partition function in eq. (13), the optimal policy in eq. (1), and the free-energy minimum in eq. (12).

A. Partition-function / KL decomposition

Writing the expectation and KL divergence explicitly and dividing by β\beta gives

Fβ[π]β=∑yπ(y)[log⁡π(y)πref(y)−r(y)β]=∑yπ(y)log⁡π(y)πref(y)er(y)/β.(A.1)\frac{\mathcal{F}_{\beta}[\pi]}{\beta} = \sum_y \pi(y)\left[\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)}-\frac{r(y)}{\beta}\right] = \sum_y \pi(y)\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)e^{r(y)/\beta}}. \tag{A.1}

The Boltzmann weight πref(y)er(y)/β\pi_{\mathrm{ref}}(y)e^{r(y)/\beta} from eq. (11) appears directly from the objective—not as an ansatz. Normalizing it by the partition function in eq. (13) gives the optimal policy in eq. (1).

Substituting πref(y)er(y)/β=Zβπβ∗(y)\pi_{\mathrm{ref}}(y)e^{r(y)/\beta}=Z_{\beta}\pi_{\beta}^{\ast}(y) then yields

Fβ[π]β=∑yπ(y)log⁡π(y)Zβπβ∗(y)=DKL(π  ∥  πβ∗)−log⁡Zβ∑yπ(y).(A.2)\frac{\mathcal{F}_{\beta}[\pi]}{\beta} = \sum_y \pi(y)\log\frac{\pi(y)}{Z_{\beta}\pi_{\beta}^{\ast}(y)} = D_{\mathrm{KL}}\left(\pi\;\|\;\pi_{\beta}^{\ast}\right) - \log Z_{\beta}\sum_y\pi(y). \tag{A.2}

Because π\pi is a normalized probability distribution, ∑yπ(y)=1\sum_y\pi(y)=1. Multiplying by β\beta therefore gives eq. (12):

Fβ[π]=−βlog⁡Zβ+βDKL(π  ∥  πβ∗).(A.3)\mathcal{F}_{\beta}[\pi] = -\beta\log Z_{\beta} + \beta D_{\mathrm{KL}}\left(\pi\;\|\;\pi_{\beta}^{\ast}\right). \tag{A.3}

Since DKL(π∥πβ∗)≥0D_{\mathrm{KL}}(\pi\|\pi_{\beta}^{\ast})\ge 0, with equality if and only if π=πβ∗\pi=\pi_{\beta}^{\ast}, the unique minimizer is eq. (1) and Fβ∗=−βlog⁡Zβ\mathcal{F}_{\beta}^{\ast}=-\beta\log Z_{\beta}.

With ϵ=−r\epsilon=-r from eq. (5), the same algebra produces the statistical-mechanics form

πβ∗(y)=πref(y)e−ϵ(y)/βZβ,Zβ=∑yπref(y)e−ϵ(y)/β.(A.4)\pi_{\beta}^{\ast}(y)=\frac{\pi_{\mathrm{ref}}(y)e^{-\epsilon(y)/\beta}}{Z_{\beta}},\qquad Z_{\beta}=\sum_y\pi_{\mathrm{ref}}(y)e^{-\epsilon(y)/\beta}. \tag{A.4}

For a uniform πref\pi_{\mathrm{ref}}, this reduces to the canonical Gibbs distribution e−ϵ/β/Zβe^{-\epsilon/\beta}/Z_{\beta}.

B. Lagrange multiplier

To minimize eq. (4) subject to ∑yπ(y)=1\sum_y\pi(y)=1, introduce a Lagrange multiplier λ\lambda:

L[π,λ]=−∑yπ(y)r(y)+β∑yπ(y)log⁡π(y)πref(y)+λ(∑yπ(y)−1).(B.1)\mathcal{L}[\pi,\lambda] = -\sum_y\pi(y)r(y) +\beta\sum_y\pi(y)\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)} +\lambda\left(\sum_y\pi(y)-1\right). \tag{B.1}

Stationarity with respect to each π(y)\pi(y) requires

∂L∂π(y)=−r(y)+β(log⁡π(y)πref(y)+1)+λ=0.(B.2)\frac{\partial\mathcal{L}}{\partial\pi(y)} = -r(y) +\beta\left(\log\frac{\pi(y)}{\pi_{\mathrm{ref}}(y)}+1\right) +\lambda =0. \tag{B.2}

Solving for π(y)\pi(y) gives

π(y)=C πref(y)er(y)/β,(B.3)\pi(y)=C\,\pi_{\mathrm{ref}}(y)e^{r(y)/\beta}, \tag{B.3}

where C=e−1−λ/βC=e^{-1-\lambda/\beta} is independent of yy. Enforcing normalization yields C−1=ZβC^{-1}=Z_{\beta} from eq. (13), and therefore the same optimal policy as eq. (1):

πβ∗(y)=πref(y)er(y)/βZβ.(B.4)\pi_{\beta}^{\ast}(y)=\frac{\pi_{\mathrm{ref}}(y)e^{r(y)/\beta}}{Z_{\beta}}. \tag{B.4}

Substituting this stationary point back into eq. (4) recovers the equilibrium free energy Fβ∗=−βlog⁡Zβ\mathcal{F}_{\beta}^{\ast}=-\beta\log Z_{\beta} from eq. (12).

  1. Rafailov et al., Direct Preference Optimization: Your Language Model Is Secretly a Reward Model, NeurIPS 2023. ↩ ↩2

  2. Schulman et al., Proximal Policy Optimization Algorithms, arXiv:1707.06347 (2017). PPO does not optimize the KL-regularized objective in eq. (2) directly: it maximizes a clipped surrogate in the ratio ρ=πθ/πθold\rho=\pi_\theta/\pi_{\theta_{\mathrm{old}}}. But at the start of each update, where πθ=πθold\pi_\theta=\pi_{\theta_{\mathrm{old}}}, the clip is inactive and ∇θ ρA=∇θlog⁡πθ A\nabla_\theta\,\rho A=\nabla_\theta \log\pi_\theta\,A—the vanilla policy gradient. With the KL-to-reference penalty folded into the reward, PPO therefore locally ascends the same KL-regularized objective; clipping only limits the step size. ↩

  3. Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, arXiv:2402.03300 (2024). ↩ ↩2