# U5. Debugging a model

> The code runs, but the model doesn’t learn. How do you find the bug?

LLM by Hand · Foundations · side trip: Under the hood · runs in your browser · last part on your computer · interactive page: https://llm.liko.page/learn/debugging/

**Boss challenge.** Fix three broken training scripts: two small NumPy ones graded in your browser, and one PyTorch script with four bugs that check.py passes only when every bug is fixed and the model gets at least 95% of the held-out points right.

Most bugs in model code don’t crash. The script runs, the loss prints, and the model just learns badly, or learns
the wrong thing. This level is a checklist for that situation. Each step comes with a small broken example.

1. **Check the shapes.**
2. **Overfit one batch.**
3. **Check the gradient** against a numeric slope.
4. **Look for the first inf or NaN.**
5. **Read the loss curve:** not falling, swinging, or falling too fast.

The boss at the end is three broken scripts. The last one runs on your computer; if you haven’t installed Python
and PyTorch yet, see the [setup page](/setup/).

All examples use the four points of level 7: (1, 3), (2, 3), (3, 7), (4, 7), fit by a line $\hat y = w x + b$ that starts at
w = 0.5, b = 0. The best line is w = 1.6, b = 1, with loss 0.8.

## 1. Check the shapes

Model code often makes the prediction a column, shape `(4, 1)`, because it is the result of `X @ W`.
The targets are often a flat list, shape `(4,)`. The loss line `((pred - y) ** 2).mean()` runs without an error.

**Question.** pred has shape (4, 1) and y has shape (4,). What is the shape of pred − y?

*Answer it on the page to check your work.*

So the loss compares every prediction with every target, 16 pairs instead of 4. The model still lowers that loss,
but the loss is asking the wrong question.

**Question.** A line fit on 4 points has a shape bug: pred − y has shape (4, 4), so the loss compares EVERY prediction with EVERY target 3, 3, 7, 7. The best the model can do is predict one number everywhere: the one with the smallest mean squared distance to 3, 3, 7, 7. Which number is it?

*Answer it on the page to check your work.*

The lab trains the line live, with the numbers of a real training run (32-bit floats, as in PyTorch).
Check the shape bug box and train: the line ends flat, and the loss stops at 4 instead of 0.8.

*[Interactive lab: Bug — open the page to use it]*

The fix is one line: make both sides the same shape, `pred[:, 0] - y` or `pred - y.reshape(-1, 1)`.
The habit that finds it: **print the shape of every array in the loss once**, and write next to it what you expect.

**If you are stuck: Why doesn’t NumPy (or PyTorch) raise an error here?**

Broadcasting is a feature: it is how `X @ W + b` adds one bias row to every row (level 1).
NumPy can’t know if you wanted a broadcast or not. So shape bugs give no error message, and only you can catch them.
`assert pred.shape == y.shape` in the loss function turns this bug into an error you can see.

## 2. Overfit one batch

Before training on everything, train on **one small batch**, for example 4 examples, many times.
A model with a few hundred weights can memorize 4 examples, so the loss should fall close to its lowest possible value.
(For the straight line that lowest value is 0.8; for an MLP it is close to 0.)

This is a test of the whole loop at once: the model, the loss, the gradients and the update.
It takes seconds, and then the amount of data cannot be the cause.

**Predict.** You train a small MLP on just 4 examples, again and again. After 500 steps the loss is still 0.69, the same as at the start. What does that tell you?

A. It needs more data
B. Something in the model, the loss or the update is broken: a working setup memorizes 4 examples easily
C. The model is too small for the task

*Answer it on the page to check your work.*

## 3. Check the gradient

Level 6 checked a hand-written gradient against a numeric one. The numeric slope at $w$ needs only the loss itself:

$$
\frac{dL}{dw} \approx \frac{L(w + h) - L(w - h)}{2h}, \quad h \text{ small}
$$

**Question.** f(w) = w², w = 3, h = 0.01. f(w + h) = 9.0601 and f(w − h) = 8.9401. What is the numeric slope (f(w + h) − f(w − h)) / (2h)?

*Answer it on the page to check your work.*

With several weights, you move one weight at a time and keep the others still. For the line fit at w = 0.5, b = 0,
the numeric gradient is dw = −21.5, db = −7.5. A correct formula gives the same two numbers.

**Predict.** Fitting y = w·x + b to (1, 3), (2, 3), (3, 7), (4, 7) at w = 0.5, b = 0. The numeric gradient is dw = −21.5, db = −7.5. Someone’s formula gives dw = −10.75, db = −3.75. What is the most likely bug?

A. The formula forgot the factor 2 from the square
B. The numeric gradient is wrong, because h is too small
C. The formula has the wrong sign

*Answer it on the page to check your work.*

Now write the numeric gradient for any number of weights. One NumPy detail: `w[i] += h` changes entry `i` of the array `w`
**in place**, so `f(w)` right after it sees the moved value. Afterwards, put `w[i]` back where it was:
the caller still needs the original `w`.

**Code question.** Write the body of the loop: the numeric slope for entry i of w. Move w[i] up by h, then down by h, and put it back where it was. Replace ____ with as many lines as you need.

Fill in the blank (`____`):

```python
def num_grad(f, w, h=1e-5):
    g = np.zeros_like(w)
    for i in range(len(w)):
        ____
    return g

x = np.array([1., 2., 3., 4.])
y = np.array([3., 3., 7., 7.])
def loss_wb(p):                       # p = [w, b]
    return ((p[0] * x + p[1] - y) ** 2).mean()

print(num_grad(loss_wb, np.array([0.5, 0.0])))   # the formula gives [-21.5, -7.5]
```

*Answer it on the page to check your work.*

**Deeper: How close is close enough?**

The numeric slope is not exact. With $h = 10^{-5}$ in 64-bit floats, a correct formula usually agrees to 6 or more
digits. Compare the **relative** difference, $|a - b| / \max(|a|, |b|)$: around $10^{-7}$ is fine, $10^{-2}$ is a bug.
In 32-bit floats the check is much less precise, so run gradient checks in 64 bits (`torch.float64`).
`torch.autograd.gradcheck` does exactly this for any PyTorch function.

## 4. Look for the first inf or NaN

A learning rate that is too large makes each step jump past the minimum and land farther away.

**Question.** f(w) = w², so the gradient is 2w. Start at w = 1 and take gradient steps w ← w − lr · 2w with lr = 1.5. What is w after 3 steps?

*Answer it on the page to check your work.*

The numbers grow until they don’t fit in a float. A 32-bit float stops at about $3.4 \times 10^{38}$; anything larger is
**inf** (infinity). In the lab, set the learning rate to 0.2 and train 200 steps: the readout shows the step where the
loss first becomes inf, and the step where it becomes **NaN** (“not a number”).

**Predict.** The loss goes 16, 86, 466, 2550, … then inf, and a few dozen steps later NaN. Where can NaN come from, if every number started out normal?

A. From inf in a calculation: inf − inf, 0 × inf and inf / inf are NaN
B. From very small numbers that round to 0
C. NaN appears at random on a GPU

*Answer it on the page to check your work.*

The other common source is a logarithm: `np.log(0)` is −inf, so a cross-entropy that takes the log of a probability that
rounded to 0 gives inf at once. That is why PyTorch’s `F.cross_entropy` takes logits and computes the log itself, safely.
Level U2 shows how floats round and why.

The habit: check `np.isfinite(loss)` (or `torch.isfinite`) every step, and stop at the **first** bad step. By the time you
notice NaN in a printout, every weight is already NaN, and the cause is hundreds of steps back.

**Try it**

In the lab, try learning rates 0.001, 0.01, 0.05 and 0.2, 200 steps each. Which one is the slowest, which one is fine,
and which one diverges? Where is the boundary between fine and diverging, roughly?

## 5. Read the loss curve

Some bugs show only in the shape of the loss curve.

**It swings up and down and never settles.** Often the gradients are never reset. `loss.backward()` **adds** into
`.grad` (level 9), so without `opt.zero_grad()` every step uses the sum of all gradients so far.

**Question.** A loop forgets `opt.zero_grad()`. For one weight, backward() gives a gradient of 2 at each of the first 3 steps. What is that weight’s .grad when opt.step() runs at step 3?

*Answer it on the page to check your work.*

In the lab, check **Forget `zero_grad`** at learning rate 0.01: the loss swings between about 1 and 17, step after step.

**It falls much faster than expected, then the model is useless.** Usually the input already contains the answer.
A **next-token model** (level 11 builds one) reads a sequence of tokens (numbers that stand for words or characters)
and, at every position, guesses the token that comes next. `<bos>` is a token that marks the start of a sequence.
The model must predict token $t + 1$ from tokens up to $t$, so the targets are the inputs **shifted by one**:
for the tokens `[<bos>, a, b, c]`, the inputs are `[<bos>, a, b]` and the targets are `[a, b, c]`.

**Predict.** A next-token model is trained on [<bos>, a, b, c]. The bug: the targets are the inputs themselves, [<bos>, a, b], instead of [a, b, c]. The training loss falls almost to 0 within a few steps. What happens when it writes text?

A. It writes well: the loss is almost 0
B. It writes nonsense: it learned to copy the token it is looking at, which needs no learning about language
C. It does not write anything

*Answer it on the page to check your work.*

A wrong causal mask does the same: if a word can see the words after it (level 15), it can read the answer.
The training loss looks wonderful; generation is nonsense.

**It doesn’t fall at all.** Check the list again from step 1: the shapes, one batch, the gradient. Then the learning rate:
try values 10 times smaller and 10 times larger.

**Deeper: The checklist, in the order to use it**

1. Print the shapes of everything in the loss. Write the shape you expect next to each one.
2. Look at one batch with your own eyes: inputs, targets, mask. Are the targets shifted? Does each label belong to its input?
3. Check the loss at step 0. For $V$ classes and random weights it should be about $\ln V$ (level 10), for example 0.69 for 2 classes.
4. Overfit one batch.
5. Check the gradients numerically, if you wrote any by hand.
6. Watch for the first inf or NaN.
7. Only then tune: the learning rate, the batch size, the model size.

## The boss: three broken scripts

Each script runs without an error. Each one learns the wrong thing. Use the checklist.

**Part A.** A line fit with three bugs. Write a working body for `train_line`.

**Code question.** Boss, part A. The teammate’s version, in the comments, has three bugs. Write a working body for `train_line` in place of ____. As many lines as you need.

Fill in the blank (`____`):

```python
def train_line(x, y, lr=0.05, steps=2000):
    # A teammate's version. It runs without an error, but the line it learns is wrong.
    #   w, b = 0.5, 0.0
    #   gw, gb = 0.0, 0.0
    #   losses = []
    #   for step in range(steps):
    #       pred = (w * x + b).reshape(-1, 1)
    #       err = pred - y
    #       losses.append((err ** 2).mean())
    #       gw += (2 * err * x.reshape(-1, 1)).mean()
    #       gb += (2 * err).mean()
    #       w = w + lr * gw
    #       b = b + lr * gb
    #   return w, b, losses
    ____

x = np.array([1., 2., 3., 4.])
y = np.array([3., 3., 7., 7.])
w, b, losses = train_line(x, y)
print(round(w, 3), round(b, 3), round(losses[-1], 3))    # should be 1.6 1.0 0.8
```

*Answer it on the page to check your work.*

**Part B.** A batch for next-token training, with two bugs. Use the collate function from level U4 and the
shift from section 5. `keep` is the mask of targets that count in the loss: `True` for real tokens, `False` for PAD.
(It is the opposite of a padding mask: `True` here means “counts”.)

**Code question.** Boss, part B. `lm_batch` builds one next-token batch: inputs X, targets Y, and keep (True where the loss should count). The teammate’s version, in the comments, has two bugs. Write a working body in place of ____.

Fill in the blank (`____`):

```python
def collate(seqs, pad=0):                 # correct (level U4)
    L = max(len(s) for s in seqs)
    P = np.full((len(seqs), L), pad)
    for i, s in enumerate(seqs):
        P[i, :len(s)] = s
    return P

def lm_batch(seqs, pad=0):
    # A teammate's version:
    #   P = collate(seqs, pad)            # (B, L), PAD at the end
    #   X = P[:, :-1]
    #   Y = P[:, :-1]
    #   keep = Y == pad
    #   return X, Y, keep
    ____

X, Y, keep = lm_batch([[1, 5, 3, 8], [1, 2, 9], [1, 7]])
print(X); print(Y); print(keep)
```

*Answer it on the page to check your work.*

**Part C, on your computer.** A PyTorch script that should learn which points lie inside a circle. It has four bugs,
one in each of `split`, `batches`, `MLP.forward` and `train_step`. Everything else is correct.

1. Download [`broken_train.py`](/files/debugging/broken_train.py) and [check.py](/files/debugging/check.py) into one folder. 
2. Copy `broken_train.py` to `my_train.py`. Run `python my_train.py` once and look at the loss and the accuracy.
3. Fix the bugs one at a time. After each fix, run `python check.py my_train.py`. It runs five checks. A check that
   fails says what it measured, not where the bug is: finding the cause is your job.
4. When all five pass, paste the last line it prints here (`U5 PASS …`). Stuck on a check? Paste the `U5 FAIL …` line
   instead: the page then gives a hint for the first failing check, a little more with each try.

The code at the end of the line only shows that you pasted it unchanged; this part relies on your honesty.

**If you are stuck: A check fails and I don’t know where to start.**

Read the numbers the check prints, then check the checklist above in order, one item at a time. For each function, print what it
returns for a few points and compare it with its docstring. Each check matches one function: split, batches, forward,
`train_step`.

The level counts as cleared once parts A and B above and this line all pass.

**Code question.** Boss, part C. Paste the last line that python check.py `my_train.py` printed, between the quotes. If a check still fails, you can paste its FAIL line for a hint on the first failing check.

Fill in the blank (`____`):

```python
line = "____"      # it looks like: U5 PASS 0.980 1a2b3c4d
print(line)
```

*Answer it on the page to check your work.*

## You can now

- Find a shape bug that gives no error message by printing shapes, before it ruins the loss.
- Check a gradient formula against the numeric slope, and overfit one batch to test the whole loop.
- Read a loss curve: diverging, NaN, stuck, swinging, or falling suspiciously fast.
