The math behind backpropagation
/AI/5 min read
Training a network means working out how much each weight contributed to the error. Backpropagation does it with one rule from calculus, applied backwards through the network.
A neural network learns by adjusting its weights. To adjust a weight you have to know two things: which direction to move it, and how much it mattered.
Backpropagation answers both. It computes, for every weight, how much the final error would change if that weight changed slightly. Then each weight steps in whichever direction makes the error smaller.
The one rule it is built on
Everything rests on the chain rule. If depends on , and depends on , then:
A network is a long chain of exactly this shape. The loss depends on the output, the output depends on the last layer, that depends on the one before it, and so on back to the input. To find how a weight near the front affects the loss at the end, you multiply the derivatives along the chain between them.
That is the whole idea. The rest is bookkeeping.
A network small enough to do by hand
Take the smallest network that still has a hidden layer: one input, one hidden neuron, one output neuron.
Each neuron does two things. First a weighted sum:
Then squashes it through a sigmoid:
And the error is measured with squared loss, halved so the derivative comes out clean:
Laid out, the chain from a weight to the loss looks like this:
The forward pass
Set some numbers. Input , target , and starting parameters:
| weight | bias | |
|---|---|---|
| hidden | ||
| output |
Push the input through:
| step | working | result |
|---|---|---|
| 0.3400 | ||
| 0.5842 | ||
| 0.0663 | ||
| 0.5166 | ||
| 0.0501 |
The network predicts 0.5166 where it should say 0.2. Now find out who is responsible.
The backward pass
Work right to left, one link at a time.
Start at the loss. How does it change with the prediction?
Through the sigmoid. Its derivative has a convenient form, , so no exponentials are needed a second time:
Multiply those two and you have how the loss responds to the output neuron's pre-activation. This combined quantity is worth naming, because everything further back reuses it:
The output weight. Since , changing changes by :
And the bias, whose derivative is just 1, so it inherits directly:
Keep going back. To reach the hidden neuron, carry through the output weight and the hidden sigmoid:
Notice the signs differ. is positive, so must come down. is negative, so must go up. Each weight gets its own direction, and they are not the same.
That reuse of is the efficiency of the whole method. The error signal is computed once at each layer and passed backwards, rather than tracing a fresh path from every weight to the loss.
Updating the weights
Each weight moves against its gradient, scaled by a learning rate. With :
| parameter | before | gradient | after |
|---|---|---|---|
| 0.6 | −0.0069 | 0.6055 | |
| −0.2 | −0.0077 | −0.1939 | |
| −0.4 | +0.0462 | −0.4369 | |
| 0.3 | +0.0791 | 0.2368 |
Run the forward pass again with those and the loss falls from 0.0501 to 0.0435, with the prediction moving from 0.5166 to 0.4951 — down towards the target of 0.2.
Repeating it
One step barely moves. Training is that step over and over:
| step | loss | prediction |
|---|---|---|
| 0 | 0.050110 | 0.5166 |
| 1 | 0.043536 | 0.4951 |
| 10 | 0.012865 | 0.3604 |
| 50 | 0.000322 | 0.2254 |
| 100 | 0.000010 | 0.2045 |
| 300 | 0.000000 | 0.2000 |
Same forward pass, same backward pass, same update — a few hundred times. Real networks differ by having millions of weights and many layers, but each weight is handled by exactly the arithmetic above.
The short version
- Backpropagation finds how much each weight contributed to the error.
- It works by applying the chain rule backwards through the network.
- The sigmoid's derivative is , so the forward pass supplies what the backward pass needs.
- The error signal is computed once per layer and reused, rather than retraced per weight.
- Each weight then steps against its own gradient, scaled by the learning rate.
- Repeat, and the loss falls.