Level 8 · Foundations · runs in your browser

Generalization

How do you train a network that also works on data it has never seen?

Level 7 made training fast and stable. A low training loss is still not the goal: the model has to work on data it has never seen. That is called generalization. This level adds two habits:

  1. A held-out set: check the model on data it never trained on.
  2. Regularization: stop the model from memorizing the noise.

1. Overfitting

A model can do well on its training data and badly on everything else. Here are 10 noisy points from the curve y = sin(πx) (a smooth wave; see the math page), and 40 more points from the same curve that the model never trains on: the held-out set.

The model is a polynomial, c0+c1x+c2x2+⋯+cdxdc_0 + c_1 x + c_2 x^2 + \dots + c_d x^d: a sum of powers of x, each times a number (a coefficient). Its degree dd is the biggest power. Each coefficient is one parameter.

Number A polynomial of degree 9, c0 + c1·x + c2·x² + … + c9·x⁹, has how many parameters?
🔒 Answer the question above to unlock
Predict firstYou fit the 10 points with degree 0, then 1, 2, … up to 9. The training loss keeps going down. What does the loss on the 40 held-out points do?
🔒 Answer the question above to unlock

Slide the degree. Each curve is the one with the smallest squared error on the training points, computed directly instead of by gradient descent, so what you see is the best each degree can do.

Too few, just right, too many parameters

Slide the degree. Watch the training loss and the held-out loss.

−101x10⁻⁴10⁻²1loss0123456789degree
Degree 3: train loss 0.0382 · held-out loss 0.1579
10 training points 40 held-out points the fit the true curve sin(πx) train loss held-out loss lowest held-out loss (degree 3)

Degree 3 has train loss 0.0382 and held-out loss 0.1579. Degree 9 passes through every training point (train loss 0.0000) and swings wildly between them: held-out loss 0.8275. The model learned the noise, not the curve. That is overfitting.

How complicated a curve a model can draw is called its capacity. More parameters give more capacity. Too little capacity misses the pattern (degree 0 or 1). Too much, with only 10 points, fits the noise (degree 9).

The same thing happens over time while training. Here is a network with 32 hidden units (97 parameters) trained on the same 10 points:

Train too long and the held-out loss rises again

Press Train and watch both losses. The ring marks where early stopping would stop.

10.10.010.0010100020003000steplosspress Train to play the run
Loss by training step, log scale.
The data and the network’s curve at step 0.
A network with 32 hidden units has 97 parameters for 10 points. Press Train.
train loss held-out loss lowest held-out so far 10 training points held-out points the network now at its best step true curve

The train loss keeps falling. The held-out loss falls, reaches its lowest point, then slowly rises again. Early stopping keeps the weights from the step where the held-out loss was lowest.

I got stuck here My training loss keeps going down, but the held-out loss went up. Is training broken?

No. Training is doing exactly what you asked: lower the loss on the training points. Past some point the only way left to lower it is to bend toward each point’s noise. That makes the curve worse everywhere else.

The training loss alone does not show that this is happening. That is the whole reason to keep a held-out set.

I got stuck here If I choose the stopping step by looking at the held-out loss, isn’t the held-out set now part of training?

A little, yes. Every decision you make by looking at it (when to stop, which degree, how much regularization) fits it a bit. So in practice there are two sets kept out of training: a validation set you use for those decisions, and a test set you look at once, at the very end, to report how good the model is.

The one rule with no exceptions: the model’s weights never train on either of them.

2. Regularization

Besides stopping early, you can make memorizing harder.

Weight decay adds a penalty for big weights to the loss: loss+λw2\text{loss} + \lambda w^2, where λ\lambda (“lambda”) is a small number such as 0.1 that sets how strong the penalty is. Its gradient, 2λw2 \lambda w, pulls every weight a little toward 0 on every step.

Number Weight decay adds a penalty to the loss: total loss = data loss + 0.1·w². The data’s gradient for w is 0 right now, w = 2, lr = 0.5. After one gradient step, what is w?
🔒 Answer the question above to unlock

In code, one step adds the penalty’s gradient to the data’s gradient, then steps downhill as usual. The same two lines work for one weight and for a whole array of weights.

CodeWrite one gradient step with weight decay: add the penalty’s gradient 2·lam·w to the data’s gradient g_data, then step downhill with lr. It must work for one weight and for an array of weights.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Go deeper Weight decay with Adam: AdamW

With plain SGD, adding λw² to the loss and shrinking every weight a little on each step are the same thing. With Adam they are not: the penalty’s gradient gets divided by v^\sqrt{\hat v} like every other gradient, so weights with big gradients are barely decayed. AdamW fixes this by skipping the loss and shrinking the weights directly after each Adam step: w←w−lr⋅λ ww \leftarrow w - \text{lr}\cdot\lambda\, w. Most language models are trained with AdamW.

🔒 Answer the question above to unlock

Here is the degree-9 polynomial again, now with weight decay on c1 … c9. Slide it up from 0.

Weight decay makes a degree-9 fit smooth

Raise the weight decay and watch the up-and-down swings and the held-out loss.

−101x10⁻⁴10⁻²1loss0123456789degree
Degree 9: train loss 0.0000 · held-out loss 0.8275
10 training points 40 held-out points the fit the true curve sin(πx) train loss held-out loss lowest held-out loss (degree 3)

A tiny decay, 0.0001, takes the held-out loss from 0.8275 down to 0.1727: the wild swings are gone. Too much, 0.1, and the curve is too flat to follow the data: train 0.1380, held-out 0.3005.

Dropout works differently. During training, each hidden unit is set to 0 at random with probability p. The network can’t depend on any single unit, so it has to store what it learns in many units. To keep the total the same size, the units that stay on are divided by 1 − p. At test time nothing is dropped. Dropping units and dividing the kept ones by 1 − p is called inverted dropout; it is the usual way to write dropout.

Number Hidden units h = [2, 4, 6, 8], dropout p = 0.5, keep mask [1, 0, 1, 0]. The kept units are divided by 1 − p. What does the third unit, 6, become?
🔒 Answer the question above to unlock

Train the 32-unit network again, with dropout p = 0.3 on its hidden units:

The same run with dropout

Train once without dropout, then check the dropout box and train again.

10.10.010.0010100020003000steplosspress Train to play the run
Loss by training step, log scale.
The data and the network’s curve at step 0.
A network with 32 hidden units has 97 parameters for 10 points. Press Train.
train loss held-out loss lowest held-out so far 10 training points held-out points the network now at its best step true curve
Try it

Train the network once without dropout and once with it. Compare the held-out loss at step 3000, and how far it climbs after its lowest point. Then go back to the polynomial and find the weight decay with the lowest held-out loss.

The GPT you write in level 21 uses the first habit: it keeps 1,500 of its 10,000 problems out of training: 500 for validation (to pick the best epoch) and 1,000 for the final test. Larger models also use dropout and weight decay.

Last, write a whole dropout layer yourself. The random part is done for you: u holds one random number between 0 and 1 per unit, and a unit is dropped when its number is below p, which happens with probability p. A comparison on an array gives one True or False per number: np.array([0.9, 0.1]) >= 0.5 is [True, False]. When you multiply, NumPy treats True as 1 and False as 0. The layer also needs to know whether it is training: at test time it must return h unchanged.

CodeWrite an inverted dropout layer. When train is False, return h unchanged. When train is True, drop each unit whose random number u is below p, and divide the kept ones by 1 − p. Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Spot overfitting: the training loss keeps falling while the held-out loss rises.
  • Compute one gradient step with weight decay, by hand and in code.
  • Apply inverted dropout to a layer with a given keep mask.

Keep in mind

  • A polynomial of degree has parameters
  • Weight decay: , so every weight gets an extra gradient
  • Dropout in training: h * keep / (1 - p); at test time nothing is dropped
  • Early stopping keeps the weights from the step with the lowest held-out loss

Common mistakes

  • Reading a low training loss as a good model: only the held-out loss shows how it does on new data.
  • Forgetting the 2 in the penalty’s gradient , or writing / 1 - p instead of / (1 - p).

Press ? for keyboard shortcuts

Reading mode · every part open, no stars