So the loss compares every prediction with every target, 16 pairs instead of 4. The model still lowers that loss, but the loss is asking the wrong question.
Debugging a model
The code runs, but the model doesn’t learn. How do you find the bug?
Side trip · best after level 9 · From NumPy to PyTorch
A short test to skip this level
Solve these 5 questions on your own. Answer all of them correctly and the level counts as cleared with three stars, and every part of the page opens. Showing an answer doesn’t count.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
train_line in place of ____. As many lines as you need.Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
lm_batch builds one next-token batch: inputs X, targets Y, and keep (True where the loss should count). The teammate’s version, in the comments, has two bugs. Write a working body in place of ____.Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
my_train.py printed, between the quotes. If a check still fails, you can paste its FAIL line for a hint on the first failing check.Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Warm-up2 questions from earlier levels
A quick review before you start. Optional. Nothing here locks the level.
Fix three broken training scripts: two small NumPy ones graded in your browser, and one PyTorch script with four bugs that check.py passes only when every bug is fixed and the model gets at least 95% of the held-out points right.
Most bugs in model code don’t crash. The script runs, the loss prints, and the model just learns badly, or learns the wrong thing. This level is a checklist for that situation. Each step comes with a small broken example.
- Check the shapes.
- Overfit one batch.
- Check the gradient against a numeric slope.
- Look for the first inf or NaN.
- Read the loss curve: not falling, swinging, or falling too fast.
The boss at the end is three broken scripts. The last one runs on your computer; if you haven’t installed Python and PyTorch yet, see the setup page.
All examples use the four points of level 7: (1, 3), (2, 3), (3, 7), (4, 7), fit by a line that starts at w = 0.5, b = 0. The best line is w = 1.6, b = 1, with loss 0.8.
1. Check the shapes
Model code often makes the prediction a column, shape (4, 1), because it is the result of X @ W.
The targets are often a flat list, shape (4,). The loss line ((pred - y) ** 2).mean() runs without an error.
The lab trains the line live, with the numbers of a real training run (32-bit floats, as in PyTorch). Check the shape bug box and train: the line ends flat, and the loss stops at 4 instead of 0.8.
Four points, one line, and some bugs
Pick a learning rate, switch a bug on or off, then train. Changing a setting starts again from w = 0.5, b = 0.
The fix is one line: make both sides the same shape, pred[:, 0] - y or pred - y.reshape(-1, 1).
The habit that finds it: print the shape of every array in the loss once, and write next to it what you expect.
I got stuck here Why doesn’t NumPy (or PyTorch) raise an error here?
Broadcasting is a feature: it is how X @ W + b adds one bias row to every row (level 1).
NumPy can’t know if you wanted a broadcast or not. So shape bugs give no error message, and only you can catch them.
assert pred.shape == y.shape in the loss function turns this bug into an error you can see.
2. Overfit one batch
Before training on everything, train on one small batch, for example 4 examples, many times. A model with a few hundred weights can memorize 4 examples, so the loss should fall close to its lowest possible value. (For the straight line that lowest value is 0.8; for an MLP it is close to 0.)
This is a test of the whole loop at once: the model, the loss, the gradients and the update. It takes seconds, and then the amount of data cannot be the cause.
3. Check the gradient
Level 6 checked a hand-written gradient against a numeric one. The numeric slope at needs only the loss itself:
With several weights, you move one weight at a time and keep the others still. For the line fit at w = 0.5, b = 0, the numeric gradient is dw = −21.5, db = −7.5. A correct formula gives the same two numbers.
Now write the numeric gradient for any number of weights. One NumPy detail: w[i] += h changes entry i of the array w
in place, so f(w) right after it sees the moved value. Afterwards, put w[i] back where it was:
the caller still needs the original w.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Go deeper How close is close enough?
The numeric slope is not exact. With in 64-bit floats, a correct formula usually agrees to 6 or more
digits. Compare the relative difference, : around is fine, is a bug.
In 32-bit floats the check is much less precise, so run gradient checks in 64 bits (torch.float64).
torch.autograd.gradcheck does exactly this for any PyTorch function.
4. Look for the first inf or NaN
A learning rate that is too large makes each step jump past the minimum and land farther away.
The numbers grow until they don’t fit in a float. A 32-bit float stops at about ; anything larger is inf (infinity). In the lab, set the learning rate to 0.2 and train 200 steps: the readout shows the step where the loss first becomes inf, and the step where it becomes NaN (“not a number”).
The other common source is a logarithm: np.log(0) is −inf, so a cross-entropy that takes the log of a probability that
rounded to 0 gives inf at once. That is why PyTorch’s F.cross_entropy takes logits and computes the log itself, safely.
Level U2 shows how floats round and why.
The habit: check np.isfinite(loss) (or torch.isfinite) every step, and stop at the first bad step. By the time you
notice NaN in a printout, every weight is already NaN, and the cause is hundreds of steps back.
In the lab, try learning rates 0.001, 0.01, 0.05 and 0.2, 200 steps each. Which one is the slowest, which one is fine, and which one diverges? Where is the boundary between fine and diverging, roughly?
5. Read the loss curve
Some bugs show only in the shape of the loss curve.
It swings up and down and never settles. Often the gradients are never reset. loss.backward() adds into
.grad (level 9), so without opt.zero_grad() every step uses the sum of all gradients so far.
opt.zero_grad(). For one weight, backward() gives a gradient of 2 at each of the first 3 steps. What is that weight’s .grad when opt.step() runs at step 3?In the lab, check Forget zero_grad at learning rate 0.01: the loss swings between about 1 and 17, step after step.
It falls much faster than expected, then the model is useless. Usually the input already contains the answer.
A next-token model (level 11 builds one) reads a sequence of tokens (numbers that stand for words or characters)
and, at every position, guesses the token that comes next. <bos> is a token that marks the start of a sequence.
The model must predict token from tokens up to , so the targets are the inputs shifted by one:
for the tokens [<bos>, a, b, c], the inputs are [<bos>, a, b] and the targets are [a, b, c].
A wrong causal mask does the same: if a word can see the words after it (level 15), it can read the answer. The training loss looks wonderful; generation is nonsense.
It doesn’t fall at all. Check the list again from step 1: the shapes, one batch, the gradient. Then the learning rate: try values 10 times smaller and 10 times larger.
Go deeper The checklist, in the order to use it
- Print the shapes of everything in the loss. Write the shape you expect next to each one.
- Look at one batch with your own eyes: inputs, targets, mask. Are the targets shifted? Does each label belong to its input?
- Check the loss at step 0. For classes and random weights it should be about (level 10), for example 0.69 for 2 classes.
- Overfit one batch.
- Check the gradients numerically, if you wrote any by hand.
- Watch for the first inf or NaN.
- Only then tune: the learning rate, the batch size, the model size.
The boss: three broken scripts
Each script runs without an error. Each one learns the wrong thing. Use the checklist.
Part A. A line fit with three bugs. Write a working body for train_line.
train_line in place of ____. As many lines as you need.Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Part B. A batch for next-token training, with two bugs. Use the collate function from level U4 and the
shift from section 5. keep is the mask of targets that count in the loss: True for real tokens, False for PAD.
(It is the opposite of a padding mask: True here means “counts”.)
lm_batch builds one next-token batch: inputs X, targets Y, and keep (True where the loss should count). The teammate’s version, in the comments, has two bugs. Write a working body in place of ____.Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Part C, on your computer. A PyTorch script that should learn which points lie inside a circle. It has four bugs,
one in each of split, batches, MLP.forward and train_step. Everything else is correct.
- Download
broken_train.pyand check.py into one folder. - Copy
broken_train.pytomy_train.py. Runpython my_train.pyonce and look at the loss and the accuracy. - Fix the bugs one at a time. After each fix, run
python check.py my_train.py. It runs five checks. A check that fails says what it measured, not where the bug is: finding the cause is your job. - When all five pass, paste the last line it prints here (
U5 PASS …). Stuck on a check? Paste theU5 FAIL …line instead: the page then gives a hint for the first failing check, a little more with each try.
The code at the end of the line only shows that you pasted it unchanged; this part relies on your honesty.
I got stuck here A check fails and I don’t know where to start.
Read the numbers the check prints, then check the checklist above in order, one item at a time. For each function, print what it
returns for a few points and compare it with its docstring. Each check matches one function: split, batches, forward,
train_step.
The level counts as cleared once parts A and B above and this line all pass.
my_train.py printed, between the quotes. If a check still fails, you can paste its FAIL line for a hint on the first failing check.Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Recap
a summary for when you finish the level
The key formulas and common mistakes appear here once you clear the level.
You can now
- Find a shape bug that gives no error message by printing shapes, before it ruins the loss.
- Check a gradient formula against the numeric slope, and overfit one batch to test the whole loop.
- Read a loss curve: diverging, NaN, stuck, swinging, or falling suspiciously fast.
Keep in mind
- A column
(n, 1)minus a flat(n,)gives an(n, n)square: no error, wrong loss - numeric slope
- inf − inf = NaN, 0 × inf = NaN, and NaN spreads to everything it touches
- next-token targets are the inputs shifted by one place: the inputs drop the last column, the targets drop the first
Common mistakes
- Changing the learning rate or the model size before checking that one batch can be memorized.
- Trusting a loss that falls very fast: the input may already contain the answer.
Press ? for keyboard shortcuts