All posts

PPO

/AI/4 min read

PPO improves a policy while refusing to move it far in one step. The clipping that enforces this is not symmetric, and the asymmetry is the part worth understanding.

A policy is whatever decides the next action — for a language model, the distribution over the next token. Reinforcement learning improves it by trying things, scoring the results, and adjusting.

The difficulty is step size. Adjust too little and training takes forever. Adjust too much and the policy lands somewhere that behaves badly, and since it is now generating its own training data, everything after that is collected from a broken policy. There is no going back to good data.

PPO is the answer that took over: improve the policy, but refuse to move it far from where it started.

The ratio

Everything is expressed through one quantity — how much more likely the new policy is to take an action than the old one was:

r=πnew(a)πold(a)r = \frac{\pi_{\text{new}}(a)}{\pi_{\text{old}}(a)}

r=1r = 1 means no change. r=1.5r = 1.5 means half again as likely. r=0.5r = 0.5 means half as likely.

Alongside it sits the advantage, AA: how much better the action turned out than expected. Positive means it beat expectations and should be made more likely; negative means the opposite.

Multiply them and you have the naive objective, rArA. Maximise it and a good action's probability is pushed up without limit — which is exactly the instability to avoid.

The clipped objective

L=min⁡(rA,  clip(r, 1−ϵ, 1+ϵ) A)L = \min\bigl(rA,\; \text{clip}(r,\, 1-\epsilon,\, 1+\epsilon)\,A\bigr)

With ϵ=0.2\epsilon = 0.2 the clip holds rr between 0.8 and 1.2. Two terms, and the smaller wins.

That min looks like a detail. It is the whole mechanism, and it does something asymmetric:

ratioA=+1A = +1clipped?A=−1A = -1clipped?
0.50.50no−0.80yes
0.80.80no−0.80no
1.01.00no−1.00no
1.21.20no−1.20no
1.51.20yes−1.50no
3.01.20yes−3.00no
THE CLIP ONLY BITES ONE WAY AT A TIME A = +1 · flat, stop pushing A = −1 · keeps falling 0.8 1.0 1.2 probability ratio L
Above 1.2 the two curves behave completely differently. That is deliberate.

Read the two curves separately.

A good action, pushed too far. Above r=1.2r = 1.2 the objective flattens. Raising the probability further gains nothing, the gradient is zero, and the update stops. This is the safety.

A bad action, made more likely. Above r=1.2r = 1.2 nothing flattens — the objective keeps falling, without limit. The gradient stays alive and the correction keeps pulling.

That asymmetry is the point of the min. A cap on how far you can reward something is prudent. A cap on how far you can correct a mistake would be dangerous. If the policy has drifted into making a bad action three times more likely, PPO does not respond by shrugging at 1.2 — it applies the full penalty and drags it back.

The same asymmetry holds mirrored on the other side: suppressing a good action too hard is not clipped either.

So the rule is not "never move more than 20%". It is never move more than 20% in the direction the data is encouraging, while leaving corrections unbounded.

Why the step limit lets you reuse data

There is a practical payoff beyond stability.

Collecting experience is expensive — for a language model, it means generating complete responses and scoring them. You would like several gradient steps out of each batch.

But the data was generated by the old policy, and after one update the current policy is different, so the data is technically off-policy and increasingly misleading. The ratio rr is exactly the measure of how stale it has become — it is 1 when the policies agree and drifts away as they diverge.

Clipping therefore does double duty: it bounds the step, and it switches off the gradient precisely when the data has become too stale to trust. That is what makes several passes over one batch safe.

What it costs

A second model. The advantage needs an expectation to compare against, which means a value model predicting how good a state is, trained alongside the policy. Roughly doubles the memory and the compute.

Hyperparameters. ϵ\epsilon, the number of passes over each batch, the value model's learning rate. PPO is more forgiving than what came before it, not forgiving.

Noisy rewards. Clipping bounds the size of a step, not its direction. If the reward signal is systematically wrong, PPO will walk steadily toward the wrong place.

In language-model alignment there is usually a further term: a penalty on drifting from the model you started with. Clipping keeps each individual step small; that penalty keeps the total distance small. They are answering different questions, and both are needed.

The short version

  • A policy decides the next action; RL improves it from scored outcomes.
  • Big steps break training irrecoverably, because the policy then generates its own bad data.
  • PPO works with the ratio of new to old probability for an action, and the advantage.
  • The objective takes the minimum of the raw and the clipped term.
  • With a positive advantage the objective flattens past 1+ϵ1+\epsilon — pushing further gains nothing.
  • With a negative advantage it does not flatten — corrections are never capped.
  • That is also what makes reusing a batch safe: the gradient dies exactly when the data goes stale.
  • The cost is a second model for the advantage, plus hyperparameters to tune.