Level 4 · Foundations · runs in your browser

Activation functions

Why does a network need a non-linear step, and which one should it use?

In level 3 one neuron drew one straight line, and no line can separate XOR. On XOR the neuron’s best cross-entropy loss was ln 2 ≈ 0.6931, the cost of a coin flip. This level stacks neurons, finds out what must sit between them, and compares the four functions used there most often.

1. Stacking neurons: does that fix it?

If one line is not enough, use several neurons side by side (a layer) and pass their outputs to another neuron. First try it without sigmoid, with plain linear layers.

Take a point x=[2,1]x = [2, 1] and two layers:

W1=[1201]W2=[11]W_1 = \begin{bmatrix} 1 & 2 \\ 0 & 1 \end{bmatrix} \qquad W_2 = \begin{bmatrix} 1 \\ 1 \end{bmatrix}

The first layer computes h=xW1h = x W_1. The second computes hW2h W_2. (In math, xW1x W_1 with nothing between the two letters means the matrix product x @ W1 in code.)

Number x = [2, 1], W1 = [[1, 2], [0, 1]] and W2 = [[1], [1]]. First h = x @ W1, then out = h @ W2. What is out?
🔒 Answer the question above to unlock

Matrix multiplication lets you group the two weight matrices first: (xW1)W2=x(W1W2)(x W_1) W_2 = x (W_1 W_2).

Number W1 = [[1, 2], [0, 1]] and W2 = [[1], [1]]. W1 @ W2 has shape (2, 1). What is its top entry?

So the two layers are really one layer with weights W1W2W_1 W_2. However many linear layers you stack, the result is still a single line. Extra layers add nothing unless something nonlinear sits between them.

Go deeper What the nonlinearity actually does

Put a nonlinear function between the layers, for example h=tanh⁡(xW1+b1)h = \tanh(x W_1 + b_1). tanh is sigmoid stretched to the range −1 to 1: tanh⁡(x)=2 sigmoid(2x)−1\tanh(x) = 2\,\text{sigmoid}(2x) - 1. Now (tanh⁡(⋅))W2(\tanh(\cdot)) W_2 can’t be rewritten as one matrix, because tanh does not distribute over addition: tanh⁡(a+b)≠tanh⁡(a)+tanh⁡(b)\tanh(a + b) \ne \tanh(a) + \tanh(b).

Each hidden neuron still draws one line, then tanh bends its output into “this side” (about +1) or “that side” (about −1). The output neuron combines those answers. Two lines plus a combination can make the XOR pattern: “between the two lines” is 1, “outside” is 0. That is the pattern you’ll see the network find in the lab below.

Any nonlinear function works in principle. The last part of this page compares the four you will meet most.

Here is a network with 6 hidden neurons, learning XOR for real in your browser. Its hidden neurons use tanh (sigmoid stretched to −1…1, see the box above). Press Train.

A 2 → 6 → 1 network learns XOR

Press Train and watch the background bend around the four points. Then turn the activation off and try again.

0110x1 →x2 ↑
loss0steps →Train to draw the loss
loss while training (last 200 records)
inputtargetoutput
[0, 0]00.000
[0, 1]10.000
[1, 0]10.000
[1, 1]00.000
steps 0 · loss (cross-entropy) — · ✗ not learned yet
background: how sure the network is of 1 target 1 target 0 a wrong answer
Predict firstGuess before you try it: if you uncheck “activation on (tanh)” and press Train, what happens? Two layers, 6 hidden neurons, but no tanh.
🔒 Answer the question above to unlock

With the activation back on, the network gets all four points right.

Try it

Press “New random weights” a few times with the activation on. The network finds a different boundary each time, but always gets all four points right. What patterns do you see? Does it ever use one line?

Each of the 6 hidden neurons in the lab uses tanh. Is tanh the best choice? The next sections compare it with three other functions.

2. Four activation functions

The function between the layers is called the activation function. Here are the four you will meet. f′(x), “f prime of x”, is the slope of f at x (level 2).

namef(x)slope f′(x)
sigmoidσ(x)=1/(1+e−x)\sigma(x) = 1 / (1 + e^{-x})σ(x) (1−σ(x))\sigma(x)\,(1 - \sigma(x))
tanhtanh⁡(x)\tanh(x)1−tanh⁡(x)21 - \tanh(x)^2
ReLUmax⁡(0,x)\max(0, x)1 if x>0x > 0, else 0
GELUx⋅Φ(x)x \cdot \Phi(x)Φ(x)+x φ(x)\Phi(x) + x\,\varphi(x)

Φ(x)\Phi(x) is the chance that a random number from the bell curve (mean 0, standard deviation 1; level 10 shows it) lands below xx, and φ\varphi is the bell curve itself. You don’t need to memorize them. What matters is the slope, because training is driven by slopes: level 6 shows that the gradient passes backwards through every layer, and at each layer it is multiplied by that layer’s f′.

Number σ(0) = 0.5. Using f′(x) = σ(x)·(1 − σ(x)), what is the slope of sigmoid at x = 0?
🔒 Answer the question above to unlock

ReLU is even simpler. Its slope has only two values.

Number ReLU(x) = max(0, x). What is its slope at x = −2?
🔒 Answer the question above to unlock

Now chain the slopes. Every layer multiplies the gradient by its own f′. (The weights multiply it too. Here we ignore them; level 7 shows how the starting weight scale changes the picture.)

ChooseA gradient flows back through 10 sigmoid layers. Each layer multiplies it by that layer’s σ′, which is at most 0.25. Start with a gradient of 1 and ignore the weights. What reaches the first layer, at most?
Number Five sigmoid layers, ignoring the weights, best case: the gradient is multiplied by 0.25 five times. 0.25⁵ = 1 / N. What is N?
🔒 Answer the question above to unlock

3. See all four at once

Drag on any panel. The solid line is f and the dashed line is its slope. The table at the bottom multiplies the slope through n layers.

Four activation functions and their slopes

Drag across any panel, or use the arrow keys, to move x. All four panels read the same x.

sigmoid
−441
f(1.0) = 0.7311 f′(1.0) = 0.1966
tanh
−441
f(1.0) = 0.7616 f′(1.0) = 0.4200
ReLU
−441
f(1.0) = 1.0000 f′(1.0) = 1.0000
GELU
−441
f(1.0) = ? f′(1.0) = ?

If every layer sits at this x, the gradient that reaches the first layer is multiplied by f′(x) once per layer: f′(x)n.

sigmoidtanhReLUGELU
f′(1.0)0.19660.42001.0000?
f′50.00030.01311.0000?
f(x) slope f′(x) y = 1 a gradient below 0.001: it has almost vanished

Three things to see:

  • Sigmoid and tanh become flat at both ends. Move x to 3: tanh′(3) is about 0.0099. A neuron whose input is that far from 0 passes almost no gradient back. This is called saturation. Stack many such layers and the gradient for the early layers shrinks toward zero: the vanishing gradient.
  • ReLU never flattens on the right. Its slope is exactly 1 for every positive input, so gradients pass through unchanged. That is why it made deep networks trainable.
  • ReLU is completely flat on the left. A neuron whose input is negative for every example gets zero gradient and never changes again.
I got stuck here If ReLU’s slope is 0 on the left, how does a neuron that went negative ever come back?

It doesn’t, through that path. If a neuron’s input zz is negative for every training example, its slope is 0 for every example, so none of its incoming weights get any gradient. Nothing moves it back. This is a dead ReLU. With a large learning rate a big update can push many neurons there at once, and a large part of the network stops learning.

Other neurons can still change what feeds into it, so it can sometimes start learning again. But this does not happen reliably. The usual ways to prevent it are a sensible starting scale for the weights (level 7), a learning rate that isn’t too large, and variants that keep a small slope on the left, such as Leaky ReLU (0.01x0.01x for x<0x < 0). GELU also has a small negative side near 0, though for very negative inputs its slope gets close to 0 too.

Number GELU(x) = x · Φ(x), and Φ(−1) = 0.1587. What is GELU(−1)? (Four decimals.)
🔒 Answer the question above to unlock
Go deeper Why GELU replaced ReLU in many Transformers

GELU looks like ReLU from far away: about 0 for very negative inputs, about xx for large positive ones. Up close it differs in two ways that help training:

  1. It is smooth. ReLU has a corner at 0, so its slope jumps from 0 to 1. GELU’s slope changes gradually, which makes the loss surface smoother for the optimizer.
  2. It is not dead on the left. Between about −3 and 0 it dips slightly below zero and still has a slope, so a neuron whose input becomes negative keeps getting gradient and can come back.

You can read GELU as “multiply xx by the chance that it should pass”: large xx passes almost fully, very negative xx is blocked, and values near 0 pass partly. The small Transformer in level 16 keeps ReLU because it is simpler to compute by hand. Many large models use GELU there instead, and level 19 shows the gated version that most models use today.

4. Write the slope yourself

Backpropagation (level 6) needs each activation’s slope as a function. Two NumPy tools make this one line. A comparison works on every element of an array at once, and .astype(float) turns True/False into 1.0/0.0:

x = np.array([-1.0, 0.0, 2.0])
x > 0                  # array([False, False,  True])
(x < 1).astype(float)  # array([1., 1., 0.])

A plain if x > 0: does not work on a whole array, because Python cannot decide if the whole array is “true”. Write ReLU’s slope for a whole array.

CodeWrite the slope of ReLU for every element of an array x (1 where x > 0, else 0).

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Now the sigmoid’s slope from the table in section 2: σ(x) (1−σ(x))\sigma(x)\,(1 - \sigma(x)). In code, σ(x)\sigma(x) is 1 / (1 + np.exp(-x)), as in the neuron of level 3, and np.exp works on every element of an array. Compute σ(x)\sigma(x) once, give it a name, then use it twice.

CodeWrite the slope of sigmoid for every element of an array x: first σ(x), then σ(x)·(1 − σ(x)). Two lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Try it

In the lab, set x = 0.5 and n = 20. Which functions still pass a usable gradient after 20 layers? Now set x = −0.5. What happens to ReLU, and what happens to GELU?

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Show that two layers with no activation between them merge into one layer.
  • Read the slope of sigmoid, tanh, ReLU and GELU, and compute them in NumPy for a whole array.
  • Explain why the gradient vanishes through many sigmoid layers.

Keep in mind

  • (x @ W1) @ W2 = x @ (W1 @ W2): without an activation, a stack is still one line
  • , at most 0.25
  • ReLU slope: 1 on the right of 0, 0 on the left
  • Through sigmoid layers, ignoring the weights, at most reaches the first layer

Common mistakes

  • Writing if x > 0: for a whole array; compare the array itself instead.
  • Computing h = W1 @ x with x as a column; in this course x is a row, so h = x @ W1.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars