Now a real stack. Each layer is h ← tanh(h @ W), 16 numbers wide, with weights drawn from a normal distribution with std 0.8/√16.
That is a little smaller than the level-7 rule of 1/√16.
Stack 30 tanh layers: how much of the signal and the gradient is left?
Drag depth and weight scale. Compare the plain stack with the residual stack.
The lines marked “plain” are the plain stack. Going up, the signal (each layer’s outputs) shrinks layer after layer. Coming back down, the gradient does the same:
in demo.py, the gradient’s std is 1.02 at the top, 0.098 ten layers down, 0.0059 twenty layers down and 0.000135 at the bottom.
The first layers get almost no learning signal, so they barely change. A deeper plain network can finally be worse than a shallow one.
I got stuck here Level 7 fixed this with a good starting scale. Why isn’t that enough?
A good starting scale keeps the numbers in range on step one. Training then changes every weight, and tanh units slowly move into their flat ends, where the slope is near 0. Thirty factors that are each a little different from 1 still multiply into something tiny or huge. The fix has to keep working while the weights change. A residual connection does.