# 8. Generalization

> How do you train a network that also works on data it has never seen?

LLM by Hand · Foundations · runs in your browser · interactive page: https://llm.liko.page/learn/generalization/

Level 7 made training fast and stable. A low training loss is still not the goal: the model has to work on
data it has never seen. That is called **generalization**. This level adds two habits:

1. **A held-out set**: check the model on data it never trained on.
2. **Regularization**: stop the model from memorizing the noise.

## 1. Overfitting

A model can do well on its training data and badly on everything else. Here are 10 noisy points from the curve
y = sin(πx) (a smooth wave; see the [math page](/math/#sin-poly)), and 40 more points from the same curve that the model never trains on: the **held-out set**.

The model is a polynomial, $c_0 + c_1 x + c_2 x^2 + \dots + c_d x^d$: a sum of powers of x, each times a number (a coefficient). Its degree $d$ is the biggest power. Each coefficient is one parameter.

**Question.** A polynomial of degree 9, c0 + c1·x + c2·x² + … + c9·x⁹, has how many parameters?

*Answer it on the page to check your work.*

**Predict.** You fit the 10 points with degree 0, then 1, 2, … up to 9. The training loss keeps going down. What does the loss on the 40 held-out points do?

A. Keeps going down too
B. Goes down, then back up
C. Goes up from the start

*Answer it on the page to check your work.*

Slide the degree. Each curve is the one with the smallest squared error on the training points, computed directly
instead of by gradient descent, so what you see is the best each degree can do.

*[Interactive lab: Capacity — open the page to use it]*

Degree 3 has train loss 0.0382 and held-out loss 0.1579. Degree 9 passes through every training point (train loss 0.0000)
and swings wildly between them: held-out loss 0.8275. The model learned the noise, not the curve. That is **overfitting**.

How complicated a curve a model can draw is called its **capacity**. More parameters give more capacity.
Too little capacity misses the pattern (degree 0 or 1). Too much, with only 10 points, fits the noise (degree 9).

The same thing happens over time while training. Here is a network with 32 hidden units (97 parameters) trained on the same 10 points:

*[Interactive lab: Curve — open the page to use it]*

The train loss keeps falling. The held-out loss falls, reaches its lowest point, then slowly rises again. **Early stopping** keeps the weights
from the step where the held-out loss was lowest.

**If you are stuck: My training loss keeps going down, but the held-out loss went up. Is training broken?**

No. Training is doing exactly what you asked: lower the loss on the training points. Past some point the only way left
to lower it is to bend toward each point's noise. That makes the curve worse everywhere else.

The training loss alone does not show that this is happening. That is the whole reason to keep a held-out set.

**If you are stuck: If I choose the stopping step by looking at the held-out loss, isn’t the held-out set now part of training?**

A little, yes. Every decision you make by looking at it (when to stop, which degree, how much regularization) fits it a bit.
So in practice there are two sets kept out of training: a **validation set** you use for those decisions,
and a **test set** you look at once, at the very end, to report how good the model is.

The one rule with no exceptions: the model's weights never train on either of them.

## 2. Regularization

Besides stopping early, you can make memorizing harder.

**Weight decay** adds a penalty for big weights to the loss: $\text{loss} + \lambda w^2$, where $\lambda$ ("lambda") is a small number such as 0.1 that sets how strong the penalty is. Its gradient, $2 \lambda w$, pulls every weight a little toward 0 on every step.

**Question.** Weight decay adds a penalty to the loss: total loss = data loss + 0.1·w². The data’s gradient for w is 0 right now, w = 2, lr = 0.5. After one gradient step, what is w?

*Answer it on the page to check your work.*

In code, one step adds the penalty’s gradient to the data’s gradient, then steps downhill as usual. The same two lines work
for one weight and for a whole array of weights.

**Code question.** Write one gradient step with weight decay: add the penalty’s gradient 2·lam·w to the data’s gradient `g_data`, then step downhill with lr. It must work for one weight and for an array of weights.

Fill in the blank (`____`):

```python
def decay_step(w, g_data, lam, lr):
    ____
    return w

print(decay_step(2.0, 0.0, 0.1, 0.5))
print(decay_step(np.array([2.0, -1.0]), np.array([1.0, 0.0]), 0.1, 0.5))
```

*Answer it on the page to check your work.*

**Deeper: Weight decay with Adam: AdamW**

With plain SGD, adding λw² to the loss and shrinking every weight a little on each step are the same thing.
With Adam they are not: the penalty's gradient gets divided by $\sqrt{\hat v}$ like every other gradient, so weights with
big gradients are barely decayed. **AdamW** fixes this by skipping the loss and shrinking the weights directly after
each Adam step: $w \leftarrow w - \text{lr}\cdot\lambda\, w$. Most language models are trained with AdamW.

Here is the degree-9 polynomial again, now with weight decay on c1 … c9. Slide it up from 0.

*[Interactive lab: Capacity — open the page to use it]*

A tiny decay, 0.0001, takes the held-out loss from 0.8275 down to 0.1727: the wild swings are gone.
Too much, 0.1, and the curve is too flat to follow the data: train 0.1380, held-out 0.3005.

**Dropout** works differently. During training, each hidden unit is set to 0 at random with probability p.
The network can't depend on any single unit, so it has to store what it learns in many units. To keep the total the same size,
the units that stay on are divided by 1 − p. At test time nothing is dropped.
Dropping units and dividing the kept ones by 1 − p is called **inverted dropout**; it is the usual way to write dropout.

**Question.** Hidden units h = [2, 4, 6, 8], dropout p = 0.5, keep mask [1, 0, 1, 0]. The kept units are divided by 1 − p. What does the third unit, 6, become?

*Answer it on the page to check your work.*

Train the 32-unit network again, with dropout p = 0.3 on its hidden units:

*[Interactive lab: Curve — open the page to use it]*

**Try it**

Train the network once without dropout and once with it. Compare the held-out loss at step 3000, and how far it climbs
after its lowest point. Then go back to the polynomial and find the weight decay with the lowest held-out loss.

The GPT you write in level 21 uses the first habit: it keeps 1,500 of its 10,000 problems out of training: 500 for validation (to pick the best epoch) and 1,000 for the final test.
Larger models also use dropout and weight decay.

Last, write a whole dropout layer yourself. The random part is done for you: `u` holds one random number between 0
and 1 per unit, and a unit is dropped when its number is below p, which happens with probability p.
A comparison on an array gives one True or False per number: `np.array([0.9, 0.1]) >= 0.5` is `[True, False]`.
When you multiply, NumPy treats True as 1 and False as 0.
The layer also needs to know whether it is training: at test time it must return h unchanged.

**Code question.** Write an inverted dropout layer. When train is False, return h unchanged. When train is True, drop each unit whose random number u is below p, and divide the kept ones by 1 − p. Several lines.

Fill in the blank (`____`):

```python
def dropout(h, u, p, train):
    # u: one random number between 0 and 1 per unit, already drawn
    ____

h = np.array([2.0, 4.0, 6.0, 8.0])
u = np.array([0.9, 0.1, 0.7, 0.3])
print("training:", dropout(h, u, 0.5, True))
print("testing: ", dropout(h, u, 0.5, False))
```

*Answer it on the page to check your work.*

## You can now

- Spot overfitting: the training loss keeps falling while the held-out loss rises.
- Compute one gradient step with weight decay, by hand and in code.
- Apply inverted dropout to a layer with a given keep mask.
