All posts

The math behind backpropagation

/AI/5 min read

Training a network means working out how much each weight contributed to the error. Backpropagation does it with one rule from calculus, applied backwards through the network.

A neural network learns by adjusting its weights. To adjust a weight you have to know two things: which direction to move it, and how much it mattered.

Backpropagation answers both. It computes, for every weight, how much the final error would change if that weight changed slightly. Then each weight steps in whichever direction makes the error smaller.

The one rule it is built on

Everything rests on the chain rule. If zz depends on yy, and yy depends on xx, then:

dzdx=dzdy×dydx\frac{dz}{dx} = \frac{dz}{dy} \times \frac{dy}{dx}

A network is a long chain of exactly this shape. The loss depends on the output, the output depends on the last layer, that depends on the one before it, and so on back to the input. To find how a weight near the front affects the loss at the end, you multiply the derivatives along the chain between them.

That is the whole idea. The rest is bookkeeping.

A network small enough to do by hand

Take the smallest network that still has a hidden layer: one input, one hidden neuron, one output neuron.

Each neuron does two things. First a weighted sum:

z=wx+bz = wx + b

Then squashes it through a sigmoid:

a=σ(z)=11+e−za = \sigma(z) = \frac{1}{1 + e^{-z}}

And the error is measured with squared loss, halved so the derivative comes out clean:

L=12(y−y^)2L = \tfrac{1}{2}(y - \hat{y})^2

Laid out, the chain from a weight to the loss looks like this:

FORWARD → w₁ zₕ aₕ zₒ aₒ L ∂L/∂aₒ ∂aₒ/∂zₒ ∂zₒ/∂aₕ ∂aₕ/∂zₕ ∂zₕ/∂w₁ ← BACKWARD: multiply the chain
Forward along the top, derivatives back along the bottom. Backprop multiplies them.

The forward pass

Set some numbers. Input x=0.9x = 0.9, target y=0.2y = 0.2, and starting parameters:

weightbias
hiddenw1=0.6w_1 = 0.6b1=−0.2b_1 = -0.2
outputw2=−0.4w_2 = -0.4b2=0.3b_2 = 0.3

Push the input through:

stepworkingresult
zhz_h0.6×0.9−0.20.6 \times 0.9 - 0.20.3400
aha_hσ(0.34)\sigma(0.34)0.5842
zoz_o−0.4×0.5842+0.3-0.4 \times 0.5842 + 0.30.0663
aoa_oσ(0.0663)\sigma(0.0663)0.5166
LL12(0.2−0.5166)2\tfrac{1}{2}(0.2 - 0.5166)^20.0501

The network predicts 0.5166 where it should say 0.2. Now find out who is responsible.

The backward pass

Work right to left, one link at a time.

Start at the loss. How does it change with the prediction?

∂L∂ao=−(y−ao)=−(0.2−0.5166)=0.3166\frac{\partial L}{\partial a_o} = -(y - a_o) = -(0.2 - 0.5166) = 0.3166

Through the sigmoid. Its derivative has a convenient form, σ′(z)=a(1−a)\sigma'(z) = a(1-a), so no exponentials are needed a second time:

∂ao∂zo=0.5166×(1−0.5166)=0.2497\frac{\partial a_o}{\partial z_o} = 0.5166 \times (1 - 0.5166) = 0.2497

Multiply those two and you have how the loss responds to the output neuron's pre-activation. This combined quantity is worth naming, because everything further back reuses it:

δo=0.3166×0.2497=0.0791\delta_o = 0.3166 \times 0.2497 = 0.0791

The output weight. Since zo=w2ah+b2z_o = w_2 a_h + b_2, changing w2w_2 changes zoz_o by aha_h:

∂L∂w2=δo×ah=0.0791×0.5842=0.0462\frac{\partial L}{\partial w_2} = \delta_o \times a_h = 0.0791 \times 0.5842 = 0.0462

And the bias, whose derivative is just 1, so it inherits δo\delta_o directly:

∂L∂b2=0.0791\frac{\partial L}{\partial b_2} = 0.0791

Keep going back. To reach the hidden neuron, carry δo\delta_o through the output weight and the hidden sigmoid:

δh=δo×w2×ah(1−ah)=0.0791×(−0.4)×0.2429=−0.0077\delta_h = \delta_o \times w_2 \times a_h(1 - a_h) = 0.0791 \times (-0.4) \times 0.2429 = -0.0077 ∂L∂w1=δh×x=−0.0077×0.9=−0.0069\frac{\partial L}{\partial w_1} = \delta_h \times x = -0.0077 \times 0.9 = -0.0069

Notice the signs differ. ∂L/∂w2\partial L/\partial w_2 is positive, so w2w_2 must come down. ∂L/∂w1\partial L/\partial w_1 is negative, so w1w_1 must go up. Each weight gets its own direction, and they are not the same.

That reuse of δo\delta_o is the efficiency of the whole method. The error signal is computed once at each layer and passed backwards, rather than tracing a fresh path from every weight to the loss.

Updating the weights

Each weight moves against its gradient, scaled by a learning rate. With η=0.8\eta = 0.8:

w←w−η∂L∂ww \leftarrow w - \eta \frac{\partial L}{\partial w}
parameterbeforegradientafter
w1w_10.6−0.00690.6055
b1b_1−0.2−0.0077−0.1939
w2w_2−0.4+0.0462−0.4369
b2b_20.3+0.07910.2368

Run the forward pass again with those and the loss falls from 0.0501 to 0.0435, with the prediction moving from 0.5166 to 0.4951 — down towards the target of 0.2.

Repeating it

One step barely moves. Training is that step over and over:

steplossprediction
00.0501100.5166
10.0435360.4951
100.0128650.3604
500.0003220.2254
1000.0000100.2045
3000.0000000.2000

Same forward pass, same backward pass, same update — a few hundred times. Real networks differ by having millions of weights and many layers, but each weight is handled by exactly the arithmetic above.

The short version

  • Backpropagation finds how much each weight contributed to the error.
  • It works by applying the chain rule backwards through the network.
  • The sigmoid's derivative is a(1−a)a(1-a), so the forward pass supplies what the backward pass needs.
  • The error signal δ\delta is computed once per layer and reused, rather than retraced per weight.
  • Each weight then steps against its own gradient, scaled by the learning rate.
  • Repeat, and the loss falls.