The next step needs your arithmetic. The lab stops at a question mark.
Backpropagation
How does the gradient flow back through the network, layer by layer?
A short test to skip this level
Solve these 5 questions on your own. Answer all of them correctly and the level counts as cleared with three stars, and every part of the page opens. Showing an answer doesn’t count.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Warm-up2 questions from earlier levels
A quick review before you start. Optional. Nothing here locks the level.
Write the backward pass for a 2→6→1 network and prove it right with a gradient check. Then train the network until all four XOR points are correct, in your browser.
In level 2 you could write the gradient by hand: one formula, two parameters. A network has thousands of parameters spread over many layers. Backpropagation is a way to find all their gradients in one pass, using a single rule again and again. That rule is the chain rule (the math page shows it with numbers).
Here is the smallest network that still has layers: one input, one neuron, one loss.
With , , , : , so and . We want .
1. The one rule
Each box knows only its own local gradient: how its output moves when its input moves. Sigmoid knows . The square knows . To get from back to , multiply the local gradients along the path. That is the chain rule.
Press “Step back” one step at a time. Each new gradient shows as “?” until you compute it in the question below.
Send the gradient back from L, one link at a time
Press “Step back” to move the gradient one box further back. The numbers in w, x, b and the target y can be edited.
- Forward: z = 0.5·2 + -1 = 0, a = σ(0) = 0.5, L = (0.5 − 1)² = 0.25
Check ∂L/∂w with a small change
Change w by ε both ways and compare the slope with the chain-rule answer. Move ε to change the size of the change.
One more multiplication reaches .
Notice what did. Once you had it, and each cost one more multiplication. Going backward means every gradient is computed once and reused by all the earlier layers. That is why it is fast.
I got stuck here Why go backward? Couldn’t I start at w and go forward?
You could, for one parameter. Starting at , you would carry “how much does z move per unit of w” forward to L. But then for you would start again from the beginning, and for every other weight again. A network with a million weights would need a million forward passes.
Going backward, you start from the one thing everyone shares, the loss, and each box passes one number to the boxes before it. One backward pass gives every gradient, for roughly the cost of one or two forward passes.
2. How do you know the gradient is right?
Go back to the definition: move up by a tiny amount , then down by the same amount, and see how much the loss changes.
This numeric gradient needs no calculus at all, just two forward passes. It is too slow to train with (two passes per weight), but it is a perfect checker. If it agrees with your backprop, your backprop is right. This is called a gradient check.
The lab writes tiny numbers the way computers do: 1e−5 means , and 2e−11 means
. This “e” has nothing to do with from level 3.
3. From numbers to matrices
A real layer does the same thing for many neurons and many examples at once. Each scalar becomes a matrix.
In the rest of this level, dz means : the gradient of the loss at a layer’s output before its
activation, one number per example and per neuron. In general it is not the error (prediction − truth) of level 2:
at a hidden layer it is whatever the chain rule gives. Only at the output, with sigmoid and cross-entropy, does it come
out as (the Deeper box below shows why).
Three rules move it around:
- gradient of a layer’s weights:
dW = X.T @ dz, whereXis the layer’s input andX.Tits transpose (level 1). “The layer’s input” means what goes into that layer: for the hidden layer it is the dataX, but for the output layer it is the hidden layer’s output (the boss below calls ita1), notX. - gradient of a layer’s bias:
db = dz.sum(axis=0), one number per neuron. A bias is like a weight whose input is always 1, so its gradient is the plain sum ofdzover the examples. (.sum(axis=0)adds down the rows, as in level 1, section 7.) - one layer back:
dz_prev = (dz @ W.T) * f'(z_prev). The@sends the gradient back through the weights. The*is element-wise: it multiplies each number by the activation’s slope at that point.
Here are the same rules on a small network with 2 inputs, 3 hidden neurons and 1 output. It uses sigmoid and the squared loss, as in section 1. Hidden neuron 1 is section 1’s neuron: , , . The output also lands on , so the backward pass starts with the same and . Step forward, then keep pressing Next to go back.
Forward with numbers, then the gradient flows back along the same lines
Press Next to step. Select a layer to see its numbers.
Layers (4)
Look at the lines into h when you step back: each one turns into ∂L/∂w, and its width now shows the gradient, not the weight.
positive weight or valuenegative weight or value A disc’s fill is its value: the stronger the color, the further from 0. forward (this step) backward: gradient “∂ −0.5” = the gradient ∂L/∂(that value). Thicker line = larger |w|. Select a layer to see its numbers.
Where does X.T @ dz come from? Look at one example first. For one input row and one output gradient,
the weight gradient is , exactly as in section 1 ().
With several examples, the loss is a sum over them, so the gradients add up. Two examples, two inputs, one output:
| example | input x | dz | x × dz (each input times dz) |
|---|---|---|---|
| 0 | [1, 2] | 0.5 | [0.5, 1.0] |
| 1 | [3, 1] | −1 | ? |
dW is the sum of the last column over both examples. X.T @ dz computes exactly that sum in one multiplication.
Our network for XOR is (4, 2) → (2, 6) → tanh → (6, 1) → sigmoid.
Rule 3 takes the gradient one layer back. Try it on the XOR network, then on two numbers by hand.
I got stuck here Why is there a transpose in dW = Xᵀ @ dz?
Let the shapes decide. is (4, 2): 4 examples, 2 inputs. The gradient at the hidden layer, dz1, is (4, 6).
The gradient must have the same shape as , which is (2, 6).
The only way to multiply a (4, 2) and a (4, 6) into a (2, 6) is (2, 4) @ dz1 (4, 6).
The 4 that disappears is the examples: the transpose adds up each weight’s gradient over all 4 examples. When a shape doesn’t line up, write the shapes down and look for the one arrangement that works.
Go deeper Why the output gradient is simply p − y
The output uses sigmoid and the loss is the cross-entropy from level 3: . Chain the two local gradients:
Multiply them and cancels: . That is why the output gradient in the code
below is just p - y (divided by 4, because the loss is the mean over 4 examples). It is also why sigmoid and
cross-entropy are almost always used together.
4. Boss: write the backward pass
The forward pass is written. You write the whole backward pass: start at the output with dz2 = (p − y) / n
(the Deeper box above shows why), then use the three rules to get all four gradients. The local gradient of tanh
is , and you already have : it is a1. In Python the square is a1 ** 2 (** is a power, as in level 2; ^ is not). The hidden tests run a gradient check on every
gradient your backward returns, with (written 1e-5).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Your gradients match the numeric ones. Now train the network. Each step moves every weight against its gradient,
exactly as in level 2. So that this cell doesn’t show the boss’s answer, it doesn’t call your backward: it gets the
gradients the slow way, by changing each number a little (the check from section 2). Your gradient check showed that
those are the same numbers your backward gives, only much slower to compute.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The lab below runs the same network, written in JavaScript instead of NumPy. Press Train and watch it learn.
A 2 → 6 → 1 network learns XOR
Press Train and watch the background bend around the four points. Then turn the activation off and try again.
| input | target | output |
|---|---|---|
| [0, 0] | 0 | 0.000 |
| [0, 1] | 1 | 0.000 |
| [1, 0] | 1 | 0.000 |
| [1, 1] | 0 | 0.000 |
In the backprop lab at the top, set the target to y = 0 and step back again. Which gradients change sign? Then set w = 5. Why does ∂L/∂w almost vanish? Look at a(1 − a) when a is close to 1.
Optional side trip: level U3 builds a small program that does this backward pass for you, the way PyTorch does, in about 50 lines of Python. Level 7 continues the main line.
Recap
a summary for when you finish the level
The key formulas and common mistakes appear here once you clear the level.
You can now
- Follow the chain rule backward through a small network by hand, multiplying the local gradients.
- Check any gradient with the numeric estimate .
- Write the backward pass of a two-layer network in NumPy and train it on XOR.
Keep in mind
- Chain rule:
dW = X.T @ dz, the same shape asWdb = dz.sum(axis=0): one number per neuron, summed over the examplesdz_prev = (dz @ W.T) * f'(z_prev); for tanh,- Sigmoid with cross-entropy: the output
dz = (p - y) / n
Common mistakes
- Writing
db = dzwithout summing over the examples, so its shape is wrong. - Using
W @ dzordz @ Wwhere the shapes call for a transpose: write the shapes down first.
- Under the hoodU3 Build your own autograd
Press ? for keyboard shortcuts