Level 7 · Foundations · runs in your browser

Optimization

How do you make training fast and stable on real data?

You can already train a network: compute the loss, compute the gradient, step downhill. Real training adds four habits that make it faster and more stable. Each one fixes a problem you can see with a few numbers:

  1. Mini-batches: don’t look at all the data before every step.
  2. Better optimizers: don’t let one steep direction set the speed for everything.
  3. Initialization: start with weights that keep the numbers a sensible size.
  4. A learning-rate schedule: take small steps at the start and at the end.

Level 8 is about the other half of training well: a model that also works on data it has never seen.

1. Mini-batches

Fit a line, y^=wx+b\hat y = w x + b, to four points: (1, 3), (2, 3), (3, 7), (4, 7). They lie roughly on y = 2x, with some noise. Start from w = 0.5, b = 0.

The errors (prediction − truth) are −2.5, −2, −5.5, −5. The gradient for w is the mean of 2 · error · x over the points.

Number Points (1, 3), (2, 3), (3, 7), (4, 7); w = 0.5, b = 0, so the errors are −2.5, −2, −5.5, −5. Using all four points, what is dL/dw = mean of 2·error·x?
🔒 Answer the question above to unlock

That used all four points. A mini-batch uses only some of them. Take a batch of just one point, (3, 7):

Number A line ŷ = w·x + b with w = 0.5, b = 0. Use a batch of just one point, (3, 7). Its error is −5.5. What is dL/dw = 2·error·x for this batch?
🔒 Answer the question above to unlock

One point pulls harder than the average: −33 instead of −21.5. Point (1, 3) alone gives only −5. Each batch has its own opinion of which way is downhill. On average the opinions agree with the full gradient, but any single step follows just one of them.

Predict firstBatch size 1, learning rate 0.02, start w = 0.5, b = 0. After each step you measure the loss on all four points. Does it go down on every step?
🔒 Answer the question above to unlock

Try it. Batch size 4 is the full data. Batch size 1 uses one point per step.

Mini-batches: each step sees only part of the data

Pick a batch size and take steps. Batch size 4 walks straight downhill; batch size 1 zig-zags.

0123-202wb
Loss over all 4 points for every (w, b): the fainter the gray, the lower the loss.
0369x=1x=2x=3x=4
The 4 data points and your line.
best possible 0.8take a step to draw the losslossstep →
Loss on all 4 points after each step (0 so far). Now: 16.375
batch size
Start: w = 0.5, b = 0. Pick a batch size and take a few steps. Each epoch reshuffles the 4 points.
your steps and your line point the last step used the best line, w = 1.6, b = 1

With batch size 4 the path is smooth. With batch size 1 it goes back and forth, and near the bottom it never settles. Even at the best line, w = 1.6 and b = 1, each point alone still pulls in its own direction.

I got stuck here If the full batch is smoother, why not always use all the data?

Because real datasets have millions of examples. One full-batch step would read all of them, just to move once. A batch of 32 gives a noisy but useful direction, for a tiny fraction of the cost, so you can take thousands of steps in the time of one full step.

The noise is not only a cost. It can move the weights away from flat places and small dips, which often helps. A common fix for the small back-and-forth moves at the end is to lower the learning rate as training goes on.

One pass through all the training data, batch after batch, is called an epoch. With 4 points and batches of 1, an epoch is 4 steps. Training usually runs for many epochs, shuffling the order each time.

Number A training set has 1,000 examples and you use mini-batches of 50. How many steps make one epoch?
I got stuck here Why shuffle the data again at the start of every epoch?

Without shuffling, every epoch has the same batches in the same order. The weights then follow the same chain of directions each time, and if the data is sorted (all 3s first, then all 4s), each part of the epoch sees only one kind of example. The last, smaller batch is also always the same few examples. A new random order each epoch gives a new set of batches, so the noise is different every time, and it helps the weights get out of flat places.

🔒 Answer the question above to unlock

2. Optimizers

Here is a loss shaped like a narrow valley:

f(w1,w2)=0.5 (w12+25 w22)f(w_1, w_2) = 0.5\,(w_1^2 + 25\,w_2^2)

Its gradient is (w1, 25w2)(w_1,\ 25 w_2). The valley is 25 times steeper across it (w2w_2) than along it (w1w_1). Start at (−4, 1). The gradient there is (−4, 25).

Plain gradient descent (SGD) does w←w−lr⋅gw \leftarrow w - \text{lr} \cdot g with lr = 0.03.

Number f(w1, w2) = 0.5·(w1² + 25·w2²), start at (−4, 1), gradient (−4, 25). One SGD step with lr = 0.03: what is the new w2?
🔒 Answer the question above to unlock

w2 jumped from 1 to 0.25, a move of 0.75. w1 moved only 0.12, from −4 to −3.88. The steep direction moves fast and the gentle one moves very slowly. You want a bigger learning rate for w1, so raise it:

ChooseThe loss is f(w1, w2) = 0.5·(w1² + 25·w2²), with gradient (w1, 25·w2). Start at (−4, 1). SGD with lr = 0.09 instead of 0.03, to make w1 move faster. What happens to w2?
🔒 Answer the question above to unlock

So with plain SGD the steepest direction sets the speed limit for every direction. Two fixes:

  • Momentum keeps a velocity uu that remembers past steps: u←βu+gu \leftarrow \beta u + g, then w←w−lr⋅uw \leftarrow w - \text{lr} \cdot u. β (beta) is usually 0.9: each step keeps 90% of the old velocity and adds the new gradient. Steps that agree (along the valley) add up. Steps that flip sign (across the valley) cancel each other.

  • Adam gives every weight its own step size. It keeps two running averages, each one mixing the newest value into the old average, for example m←0.9 m+0.1 gm \leftarrow 0.9\,m + 0.1\,g:

    • mm, the average gradient (with beta1 = 0.9),
    • vv, the average squared gradient (with beta2 = 0.999).

    It steps by lr⋅m^/v^\text{lr} \cdot \hat m / \sqrt{\hat v}. Dividing by v^\sqrt{\hat v} cancels the size of the gradient, so each weight moves about lr, whether its gradient is 4 or 25.

Here is Adam’s first step with lr = 0.3. The question asks how far w1 moves, not where it lands. m^\hat m and v^\hat v are m and v with a small correction for the first few steps. On step 1 they are simply m^=g\hat m = g and v^=g2\hat v = g^2. The square root √ always means the positive root, so (−4)2=16=4\sqrt{(-4)^2} = \sqrt{16} = 4.

Number Adam with lr = 0.3 starts at (−4, 1), gradient (−4, 25). On step 1, m̂ = g and v̂ = g², so the step is 0.3·g/√(g²). By how much does w1 change? (Give the size of the move, not the new position.)
🔒 Answer the question above to unlock

Now run all three from (−4, 1) and compare them.

Three optimizers, one valley

All three start at (−4, 1). Take steps or press Play and see which one reaches the minimum first.

minimum (0, 0)start (−4, 1)−4−2024−101w1w2

Each ring is one loss level. The rings are squeezed: the valley is 25 times steeper across (w2) than along (w1).

w1w2lossto (0, 0)
SGD-4120.54.123
last update: —
momentum-4120.54.123
last update: —
Adam-4120.54.123
last update: —
steps: 0
The gradient is g = (w1, 25·w2): 25 times steeper across the valley than along it.
SGD momentum Adam minimum start
Try it

Set momentum β to 0 and press Play. Which path does momentum follow now? Then try β = 0.99. When does a velocity that remembers more steps help, and when does it carry the weights past the bottom?

Go deeper Why Adam divides by 1 − β^t

m and v start at 0. After one step, m = 0.1·g: only a tenth of the gradient, just because it started at zero. Dividing by 1 − 0.9¹ = 0.1 undoes that, so m^=g\hat m = g. After many steps, 0.9t0.9^t is almost 0 and the correction becomes too small to matter. The same goes for v with beta2 = 0.999.

The ratio m^/v^\hat m / \sqrt{\hat v} is the key. If a weight’s gradient keeps the same sign, the ratio is near ±1 and the weight moves about lr per step. If the sign keeps flipping, m^\hat m averages toward 0 while v^\sqrt{\hat v} stays large, so the step shrinks. That is what stops the back-and-forth steps across the valley.

Write momentum’s step yourself, as two lines: first the new velocity, then the move.

CodeWrite momentum’s step: first update the velocity u from the new gradient g, then move w by lr times u. Return both.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

Now write Adam’s update. In NumPy the square root is np.sqrt, for example np.sqrt(16.0) is 4.0. Real code adds a tiny number eps (1e-8) to the bottom of the fraction (the denominator), v^+eps\sqrt{\hat v} + \text{eps}, so it never divides by 0. The function receives it as the argument eps.

CodeWrite Adam’s update. mₕₐₜ and vₕₐₜ are already computed.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

3. Initialization

Before training starts, every weight needs a value. Random is the usual answer. But how big?

To talk about “how big”, we need one word. The standard deviation (std) says how far numbers typically sit from their mean. [1, −1, 1, −1] has std 1. [10, −10, 10, −10] has std 10. (Level 10 gives the exact formula. The math page has a short example too.)

Each output of a layer is a sum over its inputs: z=x1w1+x2w2+⋯+xnwnz = x_1 w_1 + x_2 w_2 + \dots + x_n w_n. If the inputs have standard deviation 1 and the weights have standard deviation s, the sum has standard deviation about √n · s. To keep it at 1, pick s = 1/√n, where n is the number of inputs to the layer (its “fan-in”).

I got stuck here Why does the std of a sum grow like √n, not n?

Because the ups and downs partly cancel. What adds up exactly is the variance (std squared): the sum of n independent terms, each with variance s², has variance n·s². Take the square root to get back to std: √n · s. Example: 100 terms with std 0.1 each have variance 100 × 0.01 = 1, so std 1.

Number A layer has 100 inputs (fan-in n = 100). Using the rule s = 1/√n, what standard deviation should its weights have?
🔒 Answer the question above to unlock

Now stack 10 layers and get the size wrong. Each layer multiplies the standard deviation by √n · s:

Number Inputs have std 1. They pass through 10 layers with no activation, width 100, weights with std 0.2. Each layer multiplies the std by √100 · 0.2. What is the std after 10 layers?
🔒 Answer the question above to unlock

An activation function does not fix this. Try tanh:

ChooseA network has 10 layers of width 100 with tanh after every layer. The weights have std 0.01, and the inputs have std 1. Without tanh, each layer would multiply the std by √100 · 0.01. After 10 layers the outputs are…
🔒 Answer the question above to unlock

Pass 40 random inputs through 10 layers of width 100. Change the weight size and the activation.

Ten layers deep: do the layer outputs keep their size?

Pick an activation and a weight std. Each layer’s outputs should stay near std 1.

1e-121e-81e-411e4in12345678910good: about 1

The standard deviation of each layer’s outputs (layer 0 = the inputs, std 1), on a log scale from 10⁻¹² to 10⁴.

layer 1 0.495
layer 2 0.245
layer 3 0.125
layer 4 0.062
layer 5 0.031
layer 6 0.016
layer 7 8.0e-3
layer 8 3.8e-3
layer 9 1.9e-3
layer 10 9.3e-4
activation
The layer outputs have almost vanished: the last layer’s outputs are nearly all 0. Without an activation, each layer multiplies the std by about √100 × 0.05 = 0.50.
std of a layer’s outputs std = 1 on each bar vanished or grew very large the shape of each layer’s values, scaled to its own std
Go deeper Why ReLU wants √(2/n) instead of 1/√n

ReLU sets every negative number to 0. About half of the values are negative, so ReLU removes about half the variance at each layer. To correct for this, double the variance of the weights: s² = 2/n, so s = √(2/n). With n = 100 that is about 0.14. Try it in the lab with ReLU. With 0.1 the values get smaller and smaller, to about 0.02 after 10 layers. With 0.14 they stay near 0.5.

tanh is close to a straight line near 0, so it behaves like the linear case: 1/√n keeps the values from vanishing, although they still shrink, to about 0.2 after 10 layers.

4. Learning-rate schedules

Big models rarely keep one learning rate for the whole run. Two habits are common:

  1. Warmup: start with a tiny lr and raise it over the first few hundred steps.
  2. Decay: lower lr as training goes on.
Go deeper Why use warmup, and why decay?

At the start, m^\hat m and v^\hat v are averages of very few gradients, and the weights are random, so big early steps can push the weights far in a bad direction. Warmup keeps those first steps small.

Near the end, the noisy mini-batch steps go back and forth around the bottom (the small back-and-forth moves you saw in section 1). A smaller lr makes those moves smaller, so the weights end closer to the bottom.

The simplest schedule is two straight lines. Pick a highest learning rate peak, a number of warmup steps warmup, and the number of steps in the whole run, total. At step tt (counting from 0):

lr(t)={peak⋅t / warmupif t<warmuppeak⋅(total−t) / (total−warmup)otherwise\text{lr}(t) = \begin{cases} \text{peak} \cdot t \,/\, \text{warmup} & \text{if } t < \text{warmup} \\ \text{peak} \cdot (\text{total} - t) \,/\, (\text{total} - \text{warmup}) & \text{otherwise} \end{cases}

The big bracket means “use one of these two lines”. Read it as two rules: the top line while tt is less than warmup, the bottom line after that. The rate rises from 0 to peak during warmup, then falls in a straight line to 0 at the last step.

Number A schedule with peak = 0.004, warmup = 100 and total = 1000. Step 25 is still in warmup, so lr = peak · t / warmup. What is lr at step 25?
🔒 Answer the question above to unlock

After warmup, use the second line.

Number A schedule has peak = 0.004, warmup = 100, total = 1000. Step 550 is after warmup, so lr = peak · (total − t) / (total − warmup). What is lr at step 550?

Many models replace the falling straight line with a smooth curve (half a cosine wave), but the pattern is the same: up quickly, then slowly down.

The GPT you write in level 21 uses these habits: it trains with Adam on mini-batches of 64.

🔒 Answer the question above to unlock

Now put the pieces of this level together in one small training run. The loss is (w−3)2(w - 3)^2, lowest at w=3w = 3, and the schedule is given as lr_at. Write the body of the loop: each step takes the gradient, gets its learning rate from the schedule, and moves w with momentum (section 2).

CodeWrite the body of a training loop for the loss (w − 3)², with momentum and a learning-rate schedule: each step gets its own learning rate from lr_at. Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Compute one step of SGD, momentum or Adam by hand on a two-weight loss.
  • Pick a weight std for a layer from its fan-in, and say what 10 layers do to the wrong one.
  • Compute a warmup-then-decay learning rate at any step.

Keep in mind

  • One epoch = (examples / batch size) steps
  • Momentum: , then
  • Adam steps by : each weight moves about lr
  • Initialization: , = fan-in (ReLU: )
  • Every layer multiplies the std by , so 10 layers multiply it by

Common mistakes

  • Using SGD’s move lr · g for Adam: Adam divides by , so the move is about lr.
  • Adding the per-layer factors instead of multiplying them: 3 layers that each double the std give , not 6.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars