A Statistical Mechanics View of KL-Regularized RL
The math behind DPO, PPO, and the GRPO family — and whether training reaches the Gibbs minimum
Preference post-training is often formulated as KL-regularized reward maximization: increase expected reward while staying close to a reference policy. Its optimum takes the Gibbs form
where is the partition function defined in eq. (13). This same structure underlies DPO’s derivation1 and KL-regularized RLHF methods such as PPO- and GRPO-style training.2
Most explanations—including Karina Zadorozhny’s post-training guide and Ari G’s RLHF-to-DPO walkthrough—essentially stop there. They introduce , note its connection to statistical physics, observe that it is intractable but conveniently cancels from the pairwise DPO loss, derive the optimal policy in eq. (1), and move on.
But in statistical mechanics, writing down the partition function is not the end of the derivation. It is the beginning. Once is defined in eq. (13), free energy, internal energy, entropy, and temperature follow, together with a precise characterization of equilibrium. The KL-regularized RL objective admits an almost term-by-term version of the same structure.
This post works out that dictionary explicitly: reward as negative energy, as temperature, KL divergence as relative entropy, and the RL objective as a free-energy principle in eq. (4). Then I ask the question the analogy naturally suggests: if the theory predicts a unique Gibbs optimum in eq. (1), how close does actual training get to it? I test this with a GRPO-family trainer3 on GSM8K.
The KL-regularized objective
Equation (3) in Rafailov et al. (2023)1 defines the KL-regularized RLHF objective that underlies DPO. The same reward–KL tradeoff appears in PPO and GRPO; those methods optimize surrogates or sample-based estimators plus a KL divergence term for regularization with respect to reference policy. In general, for a fixed prompt , the goal is to optimize
For readability, we suppress the conditioning on and write , , and . The objective becomes
Maximizing is equivalent to minimizing
Here is a functional of the policy . Given a reference policy and reward function , the problem is to find the policy that minimizes eq. (4).
Mapping to statistical mechanics
The objective in eq. (4) has the form of a Helmholtz free energy. Each response is a state of the system, with energy
For a policy , define the mean energy and dimensionless relative entropy as
This objective is therefore
which matches term by term:
Here DPO’s plays the role of thermal energy , so its Boltzmann factor matches under the physics convention .
But shouldn’t entropy be ? Why does it involve a reference distribution, ? Ordinary entropy takes exactly this form when states have degeneracies. Suppose coarse-grained state contains microstates and has total probability . If those microstates are equally likely, each has probability , so
Let and normalize the degeneracies as . Expanding the logarithm,
The constant does not affect minimization of the free energy in eq. (4). Thus, up to an additive constant, degeneracy turns ordinary entropy into negative KL divergence relative to the normalized degeneracy measure .
Minimizing subject to (the same variational problem as in Appendix B, with replaced by ) gives the Boltzmann distribution
This is the standard form for degenerate systems; see Ellgen, Thermodynamics and Chemical Equilibrium, §21.1 and Pathria and Beale, Statistical Mechanics, §3.4.
In KL-regularized RL, is the normalized degeneracy measure . With from eq. (5) and DPO’s corresponding to thermal energy (so ),
which is the Boltzmann weight appearing in eq. (1).
The Gibbs optimum
In statistical mechanics, for fixed temperature, energy levels, and normalized degeneracy measure , the Gibbs distribution uniquely minimizes the free energy. The same variational principle holds here: once , the reward function, and the reference policy are fixed, there is a unique optimal policy that minimizes in eq. (4). In a physical system, is fixed by degeneracies; in preference optimization, is a modeling choice that may vary across setups, but within any one setup it serves as the fixed reference.
Following the steps in the appendix, the free energy decomposes as
where the partition function is
and the corresponding Gibbs policy is eq. (1), with replaced by when conditioning on is suppressed.
Since , with equality if and only if , the unique minimizer is and the equilibrium free energy is . This is exactly the free energy of a canonical ensemble. More generally, once the partition function in eq. (13) is known, equilibrium quantities such as internal energy, entropy, and heat capacity follow from it and its derivatives with respect to —the same logic summarized in the dictionary below and in eq. (8).
There are two standard ways to derive :
- Partition-function / KL decomposition. Rewrite eq. (4) so that an unnormalized Boltzmann weight appears, then normalize it by from eq. (13).
- Lagrange multiplier. Enforce the normalization constraint while setting the functional derivative of to zero; the stationary point is again eq. (1).
Full algebra for both routes is in the appendix.
The dictionary
Statistical mechanics on the left, KL-regularized RL on the right:
| Statistical mechanics | KL-regularized RL |
|---|---|
| Energy of state | (negative reward) |
| Normalized degeneracy | |
| Internal energy | |
| Entropy | |
| Temperature | DPO’s |
| Partition function | |
| Minimum free energy | |
| Free-energy gap | |
| Heat capacity |
The heat-capacity entry follows from differentiating the partition function in eq. (13). Write so that . Standard identities for the Gibbs policy in eq. (1) give and . Since implies ,
With , this is the usual canonical-ensemble relation .
Experiments on GSM8K
Everything above is exact at the level of probability distributions. To see what that ideal predicts for a real model, I ran Qwen2.5-0.5B-Instruct on GSM8K with a binary reward: when the final answer is correct and otherwise.
The Gibbs floor
For each of 64 GSM8K test prompts , I sampled 24 candidate responses from the base model, giving 1,536 completions in total. Generation used temperature 0.8, top-p 0.95, and a 512-token limit. Each candidate has a base-model log-probability and a binary correctness reward . Normalizing the base-model probabilities over the 24 candidates defines on this finite candidate set.
Treating these 24 candidates as the state space turns every expectation into a finite sum. For each prompt, eqs. (1) and (13) become
Once the candidate set has been sampled, the Gibbs policy and its expected reward, KL divergence, and free energy from eq. (4) can all be evaluated exactly on that set. Sweeping traces the optimal reward–KL frontier in Figure 1 and the corresponding free-energy landscape in Figure 2. Figure 1 also compares two ways of defining the reference weights: full-sequence probability and length-normalized probability. The remaining experiments use full-sequence probability.
GRPO-family training vs. the floor
At the reference policy the KL term is zero, so , shown as the black dotted baseline in Figure 2. If training reaches the distribution-level optimum, its free energy should fall from this baseline toward the red Gibbs minimum at from eq. (12).
To test this, I ran a GRPO-family trainer3 separately for , using group size , 6,500 GSM8K training prompts, for about one epoch (262,144 rollouts). I saved checkpoints every 32,768 rollouts and evaluated each checkpoint on the same 500 held-out prompts. The reference model starts at an expected reward of 0.348.
Why training falls short
I expected each policy to move toward its from eq. (1) and each free-energy curve to approach its dashed target from eq. (12). Figure 3 shows that this was too optimistic: after one epoch, none of the runs reached its Gibbs minimum. Reward improved, but not enough to offset the accompanying -weighted KL cost, so the measured free energy stayed above the target.
Likely explanations:
- Limited expressivity. Optimization is restricted to the parameter space of a 0.5B model, which may not contain the true Gibbs policy.
- Non-convex loss. Even within that parameter space, the loss is non-convex, so gradient descent is not guaranteed to find the global minimum.
Takeaways
- KL-regularized RL is free-energy minimization in statistical physics: eq. (4) is the RL free energy and eq. (8) is the dictionary.
- Given , reference policy , and reward function , there is a unique policy in eq. (1) that minimizes the free energy.
- From in eq. (13) we can read off internal energy, entropy, temperature, free energy, and heat capacity.
- Training may not reach that distribution-level optimum.
Code and full experiments: github.com/wang-zhongwei/stat-mech-dpo
Appendix: two derivations of the optimal policy
Both routes start from eq. (4), equivalently written as
and the normalization constraint . They recover the partition function in eq. (13), the optimal policy in eq. (1), and the free-energy minimum in eq. (12).
A. Partition-function / KL decomposition
Writing the expectation and KL divergence explicitly and dividing by gives
The Boltzmann weight from eq. (11) appears directly from the objective—not as an ansatz. Normalizing it by the partition function in eq. (13) gives the optimal policy in eq. (1).
Substituting then yields
Because is a normalized probability distribution, . Multiplying by therefore gives eq. (12):
Since , with equality if and only if , the unique minimizer is eq. (1) and .
With from eq. (5), the same algebra produces the statistical-mechanics form
For a uniform , this reduces to the canonical Gibbs distribution .
B. Lagrange multiplier
To minimize eq. (4) subject to , introduce a Lagrange multiplier :
Stationarity with respect to each requires
Solving for gives
where is independent of . Enforcing normalization yields from eq. (13), and therefore the same optimal policy as eq. (1):
Substituting this stationary point back into eq. (4) recovers the equilibrium free energy from eq. (12).
-
Rafailov et al., Direct Preference Optimization: Your Language Model Is Secretly a Reward Model, NeurIPS 2023. ↩ ↩2
-
Schulman et al., Proximal Policy Optimization Algorithms, arXiv:1707.06347 (2017). PPO does not optimize the KL-regularized objective in eq. (2) directly: it maximizes a clipped surrogate in the ratio . But at the start of each update, where , the clip is inactive and —the vanilla policy gradient. With the KL-to-reference penalty folded into the reward, PPO therefore locally ascends the same KL-regularized objective; clipping only limits the step size. ↩
-
Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, arXiv:2402.03300 (2024). ↩ ↩2