All posts

The math behind cross-entropy loss

/AI/5 min read

Cross-entropy turns a set of predicted probabilities into one number saying how wrong they were. It is built so that being confidently wrong costs far more than being unsure.

A classifier does not output an answer. It outputs a probability for every possible answer.

Cross-entropy is how those probabilities get scored. It looks at the probability the model gave to the option that was actually correct, and turns it into a single number — small when that probability was high, large when it was low.

The formula

CE=−∑iyilog⁡(pi)\text{CE} = -\sum_i y_i \log(p_i)

pip_i is the probability the model gave to option ii. yiy_i is 1 for the correct option and 0 for all the others.

Those zeros do most of the work. Every term where yi=0y_i = 0 multiplies out to nothing, so the whole sum collapses to a single term:

CE=−log⁡(pcorrect)\text{CE} = -\log(p_{\text{correct}})

The probabilities the model assigned to the wrong options never appear. Only one number matters: how much belief was placed on the right answer.

Why the log

The log is what makes the penalty non-linear.

probability on the right answer−log⁡(p)-\log(p)
0.950.05
0.50.69
0.053.00
0.0055.30

Notice the shape. Between 0.95 and 0.5 the loss rises by 0.64. Between 0.05 and 0.005 — the same tenfold drop in probability — it rises by 2.30 again, and it keeps rising without limit as the probability approaches zero.

A model that is merely unsure gets a mild penalty. A model that is confident and wrong gets an enormous one. That asymmetry is deliberate: it is far worse to rule out the truth than to hesitate about it.

The minus sign is bookkeeping. Probabilities are at most 1, so their logs are at most 0. Negating makes the loss positive, with 0 as perfection.

Working one through

A model sorts a support ticket into one of four buckets. The ticket is really a bug report.

The model gives raw scores — logits — and softmax turns them into probabilities. Start with a model that is leaning the right way but not certain, with logits [1.2, 2.8, 0.4, -0.6]:

bucketprobability
billing0.1523
bug0.7542
feature0.0684
spam0.0252

Those sum to 1.0000. The loss is −log⁡(0.7542)=0.2822-\log(0.7542) = 0.2822.

Now sharpen the same model's opinion, so it puts 0.9696 on bug. The loss falls to 0.0309.

And now break it, so it puts 0.9608 on billing and only 0.0194 on bug. The loss jumps to 3.9399.

LOSS AGAINST BELIEF IN THE RIGHT ANSWER 0 0.5 1 probability given to the right answer 3.94 confident and wrong 0.28 0.03 loss
The same curve throughout. Sliding left costs far more than sliding right saves.

Between the second and third case the model's belief in the truth fell by a factor of about fifty, and the loss rose by a factor of about a hundred and thirty. That gap is the point of the whole function.

The two-class version

With only two options, one probability determines the other — if the chance of yes is pp, the chance of no is 1−p1 - p. So the formula is usually written as:

BCE=−[ ylog⁡(p)+(1−y)log⁡(1−p) ]\text{BCE} = -\left[\, y \log(p) + (1 - y)\log(1 - p) \,\right]

It looks like two terms but only ever one runs. When the answer is yes, y=1y = 1 kills the second half and the loss is −log⁡(p)-\log(p). When it is no, y=0y = 0 kills the first half and the loss is −log⁡(1−p)-\log(1-p).

With the true answer being yes: a prediction of 0.85 costs 0.1625, and a prediction of 0.30 costs 1.2040.

In a language model

A language model is a classifier whose options are every token in its vocabulary — often more than 50,000 of them.

The scoring is unchanged. At each position, find the probability it gave to the token that actually came next, take the negative log, and average over the sequence:

L=1N∑t=1N−log⁡(pt(actual token))L = \frac{1}{N}\sum_{t=1}^{N} -\log\bigl(p_t(\text{actual token})\bigr)

That average is usually reported after exponentiating it, as perplexity:

perplexity=eL\text{perplexity} = e^{L}

An average loss of 2.10 becomes a perplexity of 8.17. The exponential undoes the log, which puts the number back on a scale you can picture: roughly how many options the model was effectively torn between at each step. A perplexity of 8 means it was about as uncertain as someone guessing between eight equally likely tokens. Lower is better, and 1 would mean it was certain every time.

Why the gradient is so simple

Cross-entropy is almost always applied straight to the output of a softmax, and that pairing produces something unusually clean. The derivative of the loss with respect to a raw logit is:

∂CE∂zi=pi−yi\frac{\partial \text{CE}}{\partial z_i} = p_i - y_i

Predicted minus actual. Nothing else.

For the first case above the four gradients are:

bucketpi−yip_i - y_i
billing+0.1523
bug−0.2458
feature+0.0684
spam+0.0252

The correct class gets a negative gradient, so its logit is pushed up. Every wrong class gets a positive one, so their logits are pushed down, and by exactly the amount of belief each was wrongly given. They sum to zero, which is what keeps the probabilities summing to 1.

There is no division, no chain of intermediate terms, nothing that can blow up. That is why frameworks fuse softmax and cross-entropy into one operation instead of computing them separately.

The short version

  • Cross-entropy scores a set of predicted probabilities against the answer that was actually right.
  • The zeros in the one-hot label collapse the sum to −log⁡(pcorrect)-\log(p_{\text{correct}}) — only belief in the truth counts.
  • The log makes the penalty grow without limit as that belief approaches zero, so confident mistakes are punished hardest.
  • The two-class version is the same formula with one of its two halves always switched off.
  • Language models use it per token and report eLe^{L} as perplexity.
  • Paired with softmax, its gradient is just predicted minus actual, which is why the two are computed together.