Level 5 · Foundations · runs in your browser

Multilayer perceptron

What do you gain by stacking layers of neurons?

In level 3 one neuron drew one straight line. In level 4, XOR needed a second neuron and a bend between the layers. That small network has a name: a multilayer perceptron, or MLP. It is just layers of neurons, each layer feeding the next, with an activation function in between.

This level asks what you get from more neurons per layer (width) and more layers (depth). First, the forward pass written the way every library writes it: as matrices.

1. One hidden layer, as matrices

Put the inputs in a matrix XX with one row per example. A hidden layer of nn neurons is one matrix multiplication, a bias, and an activation. The output layer is one more:

H=act(XW1+b1)p=σ(HW2+b2)H = \text{act}(X W_1 + b_1) \qquad p = \sigma(H W_2 + b_2)

As in level 4, XW1X W_1 is the matrix product X @ W1 in code. Column jj of W1W_1 holds neuron jj‘s two weights, so W1W_1 is (2, n). HH holds the hidden values: one row per example, one column per hidden neuron.

Try one example by hand, with ReLU as the activation:

x=[2,1]W1=[1−101]b1=[0,0]W2=[2−1]b2=0.5x = [2, 1] \quad W_1 = \begin{bmatrix} 1 & -1 \\ 0 & 1 \end{bmatrix} \quad b_1 = [0, 0] \quad W_2 = \begin{bmatrix} 2 \\ -1 \end{bmatrix} \quad b_2 = 0.5
Number x = [2, 1], W1 = [[1, −1], [0, 1]], b1 = [0, 0], activation ReLU. The hidden layer is h = ReLU(x @ W1 + b1). What is h[1], the second hidden value?
🔒 Answer the question above to unlock

Now the output layer, then the shapes.

Number h = [2, 0], W2 = [[2], [−1]], b2 = 0.5. What is z = h @ W2 + b2, the number that goes into the output sigmoid?
Shape X holds 50 examples with 2 numbers each, (50, 2). W1 is (2, 16). What shape is H = tanh(X @ W1 + b1)?
🔒 Answer the question above to unlock

2. A harder pattern: two spirals

XOR had four points. Here are 200, in two arms that wind around each other. No straight line separates them, and neither do two or three. Before you press anything:

Predict firstA network with 1 hidden layer of only 2 tanh units trains on the two spirals. Guess before you try it: how does it do?
🔒 Answer the question above to unlock

The lab below starts with 1 hidden layer of 8 tanh units. Press Train. The Map view shows the network’s answer as a background color, one color for each arm; switch to 3D to see the same answer as a height over the plane.

A network learns two spirals

Pick the size of the network, then press Train. The background shows the network’s answer at every point.

x₁ → x₂ ↑ −1 −10 01 1
Color is the network’s answer at every point; the stronger the color, the surer the network.
loss0steps →Train to draw the loss
X (200, 2),@ W1 (2, 8) → (200, 8),@ W2 (8, 1) → (200, 1)
hidden layers
units per layer
activation
step 0 · loss – · 0% of 200 correct · parameters 33
arm 1 arm 2 background: the network’s answer, stronger = surer

Now change the width. Set 2 units per layer and train again. Then try 32. With 2 units the network can only bend the plane twice, and it fails on most of the spiral. With enough units the background starts to follow the arms.

I got stuck here Why does it find a different answer every time I press New random start?

Every run starts from different random weights, and training only ever walks downhill from where it starts. So two runs can end in different valleys: different boundaries, sometimes one that fits and one that gets stuck. That randomness is normal. Real training fixes the random seed when it needs to repeat a run exactly.

I got stuck here How does the lab know which way to move all those weights?

It computes the gradient of the loss for every weight, then takes a small step against it, exactly the gradient descent from level 2, with a smarter step size you will meet in level 7. How it gets a gradient for a weight buried inside a hidden layer is the chain rule, applied layer by layer. That is level 6.

3. How big is it?

2 inputs3 hidden neurons1 output
A 2 → 3 → 1 network: 6 lines into the hidden layer, 3 into the output. Each line is one weight.

Every line between two neurons is one weight, and every neuron after the input has one bias.

Number A network 2 → 8 → 1: 2 inputs, one hidden layer of 8 units, 1 output. How many weights and biases does it have in total?
🔒 Answer the question above to unlock

Now add a layer.

Number A network 2 → 8 → 8 → 1: 2 inputs, two hidden layers of 8 units, 1 output. How many weights and biases?

The parameter count in the lab is now unlocked. Compare your answers with it.

Here is the same idea in 3D. It starts with the example from section 1 (x=[2,1]x = [2, 1], ReLU). Then pick a width and a depth and count the lines: each one is a weight.

Every line is one weight

Pick a width and a depth. Press Next to run one example through it.

xHpW1W2
Press Next to start. · 9 parameters in all
Layers (3)

Look at the lines between two hidden layers: width × width of them. Width 6 has four times as many as width 3.

positive weight or valuenegative weight or value A disc’s fill is its value: the stronger the color, the further from 0. forward (this step) Thicker line = larger |w|. Select a layer to see its numbers.

4. What each hidden unit does

Check show each first-layer unit’s line. Every unit in the first layer is a neuron from level 3: it draws one straight line and reports which side a point is on. The layer after it combines those reports. Many lines, combined with bends in between, can trace almost any shape: here, a spiral.

Go deeper Width or depth?

A wide single hidden layer can, in principle, approximate any reasonable function: add enough units and their lines can split the plane into as many pieces as you like. But “enough” can be enormous.

Depth reuses work. A second layer does not see points, it sees the first layer’s answers, so it can combine pieces that the first layer has already found. That is why the same number of parameters often does more when it is spread over two or three layers than when it is all in one. Try it below.

Depth has a cost: the gradient has to travel back through every layer, and level 4 showed how it can get smaller at every layer. Levels 6 and 7 deal with that.

Try it

Compare 1 layer of 32 units with 3 layers of 8 units. Which has more parameters? Which one draws the spiral better after 1000 steps? Then switch to ReLU: the boundary is made of straight pieces instead of curves. Why?

5. Write it yourself

Write the hidden layer of a one-hidden-layer MLP with tanh. Everything is a matrix, so one line handles a whole batch of examples. In NumPy, tanh is np.tanh: np.tanh(np.array([0., 1.])) gives [0., 0.76], one tanh per element.

CodeWrite the hidden layer: tanh of X @ W1 plus b1.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Now add depth. A second hidden layer is the same line again, with H1H_1 as its input instead of XX. Use ReLU this time: in NumPy it is np.maximum(0, …), as in level 4. This is the 2 → 2 → 2 → 1 version of the network in section 1:

H1=ReLU(XW1+b1)H2=ReLU(H1W2+b2)out=H2W3+b3H_1 = \text{ReLU}(X W_1 + b_1) \qquad H_2 = \text{ReLU}(H_1 W_2 + b_2) \qquad \text{out} = H_2 W_3 + b_3
CodeWrite the forward pass of a network with two hidden ReLU layers, then return the output of the last layer (W3 and b3, no sigmoid). Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Compute an MLP’s forward pass by hand for one example: hidden layer, activation, output.
  • Give the shape of every layer’s output for a batch of examples.
  • Count the weights and biases of a network from its layer sizes.

Keep in mind

  • ,
  • (N, d_in) @ (d_in, n) → (N, n): one row per example, one column per hidden unit
  • Parameters per layer: inputs × outputs weights + one bias per output
  • Each first-layer unit draws one line; later layers combine those lines

Common mistakes

  • Counting only the lines (weights) and forgetting one bias per neuron.
  • Adding the bias after the activation instead of before it: act(X @ W1 + b1).

Press ? for keyboard shortcuts

Reading mode · every part open, no stars