# 21. Write your own GPT

> Can you write a working model from a starter file that has the names but no code?

LLM by Hand · Theory · runs on your computer · interactive page: https://llm.liko.page/learn/write-a-gpt/

**Boss challenge.** A decoder-only GPT with pre-norm blocks and RMSNorm learns two-digit addition. Your code must pass a fixed-weights check, and your trained model must get at least 950 of the 1,000 unseen problems fully right.

Until now every level gave you most of the code. This one gives you only the task and a **starter file**: a Python file
in which every class and function already has its name and a one-line description, but no code inside.
You can reread earlier levels as often as you like, but you write every line yourself, on your own computer.
If you haven’t run Python and PyTorch locally yet, start with the [setup page](/setup/).
Recommended before this boss: levels U4 (feeding data to a model) and U5 (debugging a model). They cover the batches,
padding and bugs you will meet here. If you skipped them, this is the batch loop you need for one epoch:

```python
# X, Y: every training x and y. perm: a new order every epoch.
perm = torch.randperm(len(X))
for i in range(0, len(X), 64):
    idx = perm[i:i + 64]               # the next 64 ids
    xb, yb = X[idx], Y[idx]
    # forward, loss, zero_grad, backward, step
```

## The task

- **What the model does.** It reads `<bos>23+58=` and continues with `81<eos>`. One character per token.
  `<bos>` is the start token: every sequence begins with it, so the model has an input before the first digit.
  `<eos>` is the end token: the model writes it when the answer is complete.
  This is exactly how a chat model writes: it predicts the next token, again and again.
- **Data.** Every pair a, b from 0 to 99: 10,000 problems. Shuffle them. Train on 8,500. Use 500 for validation, to
  pick the best epoch. Keep the last 1,000 apart: the model never sees them, and you make no choice based on them.
- **Vocabulary.** 15 tokens: `<pad>`, `<bos>`, `<eos>`, the ten digits, `+` and `=`.
- **Model.** Decoder only: embeddings, a few blocks of causal self-attention and FFN, then an output layer.
  No encoder and no cross-attention. The blocks are **pre-norm** (level 16) and use **RMSNorm** (level 19), with
  learned position embeddings.
- **Pass.** Two checks. First, your code loads fixed weights and must compute the same logits as the reference.
  Second, your trained model writes the answers to the 1,000 unseen problems, and at least 950 must be fully right.

Do the steps in order. Each step has a check that tells you a piece works before you use it in the next step.

## Step 0: your tools

Download these files into one folder:

1. [skeleton.py](/files/write-a-gpt/skeleton.py), the starter file: names, shapes and one-line descriptions, no hints. Copy it to `my_gpt.py` and complete every `# TODO`. 
   Keep the class and attribute names: the boss check loads weights by those names.
   If you are stuck, [`skeleton_hints.py`](/files/write-a-gpt/skeleton_hints.py) is the same file with comments that say which lines to write. Try without it first. 
2. [check.py](/files/write-a-gpt/check.py) and [`check_weights.json`](/files/write-a-gpt/check_weights.json), the boss check.
   You run it on your own file at the end of step 2 and step 5.
3. The **translation table** in level 9, [section 7](/learn/numpy-to-pytorch/#7-a-translation-table-to-keep).
   Almost every line you write is a NumPy line from an earlier level, translated: `nn.Embedding`, `view`, `transpose`,
   `masked_fill`, `F.cross_entropy` with `ignore_index`, and the epoch loop. One name differs from NumPy:
   `np.triu(…, k=1)` is `torch.triu(…, diagonal=1)` in PyTorch.

**Deeper: What the starter file contains**

```python
PAD, BOS, EOS = 0, 1, 2
ITOS = ["<pad>", "<bos>", "<eos>"] + list("0123456789") + ["+", "="]   # 15 tokens, ids 0–14
STOI = {c: i for i, c in enumerate(ITOS)}                             # {'<pad>': 0, '<bos>': 1, …}
MAXLEN = 11          # the longest sequence: <bos> 9 9 + 9 9 = 1 9 8 <eos>
IGNORE = -100        # F.cross_entropy(..., ignore_index=IGNORE) skips these targets

def encode(a, b): ...          # (23, 58) → 11 ids, padded with <pad>
def make_xy(ids): ...          # x, y shifted by one; y = IGNORE before the answer and on <pad>
def split(seed=0): ...         # given: 8,500 train, 500 validation, 1,000 unseen (the same split check.py uses)

class RMSNorm(nn.Module): ...  # divide by the root mean square, times a learned gain
class Attention(nn.Module): ...# q, k, v → heads → scores → causal mask → softmax → join heads → output
class Block(nn.Module): ...    # pre-norm: x = x + attn(norm1(x)); x = x + ffn(norm2(x))
class GPT(nn.Module): ...      # embeddings → blocks → final norm → output layer: (B, L) ids → (B, L, 15)

def overfit_one_batch(): ...   # step 3
def train(epochs): ...         # step 4, ends by saving model.pt for check.py
def solve(model, a, b): ...    # step 5: greedy until <eos>
```

**Deeper: Lines you copy, choices you make**

Some lines are the same in almost every PyTorch script. You can copy them without a reason of your own:
`torch.manual_seed(0)`, `opt.zero_grad(); loss.backward(); opt.step()`, Adam with learning rate 0.001, and batches of 64.

Other lines are choices that change what the model learns. You should be able to say why you made each of them:

- which positions count in the loss;
- how many blocks, and how wide they are;
- pre-norm or post-norm;
- how positions enter the model.

The encoder–decoder of side trip [N7](/learn/encoder-decoder/) uses `label_smoothing=0.1`, as the 2017 design did.
It is a choice, not a habit. This task doesn’t need it.

## Step 1: the data, shifted by one

Every token has an id. These are the ids the checks on this page use:

| token | `<pad>` | `<bos>` | `<eos>` | `0` … `9` | `+` | `=` |
|---|---|---|---|---|---|---|
| id | 0 | 1 | 2 | 3 … 12 | 13 | 14 |

So the digit 2 has id 5 and the digit 8 has id 11. Take the problem 23 + 58. As one sequence of tokens it is
`<bos>23+58=81<eos>`.

**Question.** How many tokens are in <bos>23+58=81<eos>?

*Answer it on the page to check your work.*

A batch is a rectangle: every sequence in it must have the same length. So every sequence gets `<pad>` tokens at the
end, up to the length of the longest possible one.

**Question.** Sequences in a batch must all have the same length. The longest problem is 99 + 99. How many tokens are in <bos>99+99=198<eos>?

*Answer it on the page to check your work.*

The input x is the padded sequence without its last token, and the target y is the same sequence without its first token.
So x and y have 10 tokens each, and y[t] is the token that comes after x[t].

**Question.** Positions count from 0: x = <bos> 2 3 + 5 8 = 8 1 … At which position t does the model make its guess for the first digit of the answer?

*Answer it on the page to check your work.*

The model should be graded only on the answer. The digits of a and b are random, so no model can predict them, and
`<pad>` is not something anyone wants it to say. So in y, every target before the answer and every `<pad>` is replaced
by a special value, `IGNORE = -100`, and `F.cross_entropy(..., ignore_index=-100)` skips those positions. This page
calls it the **loss on the answer digits only**.

**Question.** y for 23 + 58 is 2 3 + 5 8 = 8 1 <eos> <pad>. With the loss on the answer digits only, every target before the answer and every <pad> becomes IGNORE. How many targets still count?

*Answer it on the page to check your work.*

So “the answer digits only” always includes `<eos>`: the model is graded on the answer and on knowing when to stop.

Why is there no padding mask, as in level 15? Here every `<pad>` comes at the end, after `<eos>`, and the causal mask
already stops each position from looking at later positions. So no real token ever looks at a `<pad>`. The `<pad>`
positions do look at the tokens before them, but nothing uses their outputs: their targets are IGNORE, and generation
stops at `<eos>`.

Now check your answers in the lab, then try a few problems of your own. Compare x and y position by position, and see which targets the loss skips.

*[Interactive lab: Add data — open the page to use it]*

**Check:** print three examples as characters, as ids, as x and as y. Check that every y is its x moved one step left.

**Code question.** Split one sequence of ids into the input x and the target y.

Fill in the blank (`____`):

```python
def make_xy(ids):
    x, y = ____
    return x, y

ids = [1, 5, 6, 13, 8, 11, 14, 11, 4, 2, 0]   # <bos> 2 3 + 5 8 = 8 1 <eos> <pad>
x, y = make_xy(ids)
print("x:", x)
print("y:", y)
```

*Answer it on the page to check your work.*

**If you are stuck: I understand the shift, but why remove the first and the last token?**

x and y come from the same sequence. x drops the last token, and y drops the first. So at every position t,
y[t] is “the token after x[t]”. One pass through the model makes ten guesses at once, one per position,
and the causal mask makes sure no guess can see its own answer.

## Step 2: the forward shape

Before you write the model, watch a tiny trained GPT read `<bos>2+3=` and write the next token. It has the same parts as
yours, only with very small sizes, so you can see every number. Step through it and note the shape after each part.

*[Interactive lab: Microscope — open the page to use it]*

Now build your own model with random weights and pass one batch through it. Nothing is trained yet. Only the shapes matter.

One difference from levels 16 and 17: the input is just `tok(ids) + pos(positions)`, with **no** × √`d_model`.
Those levels use the factor because their position table is fixed (numbers between −1 and 1), and the token numbers
must not be hidden by it. Here both tables are learned `nn.Embedding` layers that start at the same size, so
training sets the balance between them. The boss check expects no factor.

**Question.** A batch of 64 examples, each padded to 11 tokens, so each x has 10. The vocabulary has 15 tokens. What shape are the logits?

*Answer it on the page to check your work.*

The hardest line in the model splits the heads. In level 15 each head took its own slice of columns.
In code, the last axis of size `d_model` is reshaped into h heads of `d_k` numbers. Then the heads move in front of the
length, so every head gets its own (L, `d_k`) table:

```python
q = q.view(B, L, h, d_k).transpose(1, 2)    # (B, L, d_model) → (B, h, L, d_k)
```

**Question.** q has shape (32, 10, 64): batch 32, length 10, `d_model` 64. With 4 heads, dₖ = 16. After view(32, 10, 4, 16) and transpose(1, 2), what shape is q?

*Answer it on the page to check your work.*

In NumPy the same split is a reshape and a transpose with all four axes in their new order ([level 9’s table](/learn/numpy-to-pytorch/#7-a-translation-table-to-keep)):

**Code question.** Split the last axis into h heads and move the heads in front of the length: (B, L, d) → (B, h, L, d // h).

Fill in the blank (`____`):

```python
def split_heads(x, h):
    B, L, d = x.shape
    return ____

x = np.arange(8.0).reshape(1, 2, 4)   # 1 example, 2 tokens, d = 4
print(split_heads(x, 2))
```

*Answer it on the page to check your work.*

After the attention, the heads go back together: move the length in front of the heads again, then join the last two axes.

```python
y = out.transpose(1, 2).contiguous().view(B, L, d_model)    # (B, h, L, d_k) → (B, L, d_model)
```

**Question.** After the attention, out has shape (32, 4, 10, 16): batch 32, 4 heads, length 10, dₖ 16. What shape is out.transpose(1, 2).contiguous().view(32, 10, -1)?

*Answer it on the page to check your work.*

**If you are stuck: view says “view size is not compatible with input tensor’s size and stride”.**

`view` only relabels the numbers as they lie in memory. It never moves them. After `transpose`, the numbers that should
be next to each other are not next to each other in memory, so `view` refuses. `.contiguous()` makes a copy with
the numbers in the new order, and then `view` works. `.reshape(B, L, d_model)` does both steps for you.
The error says: `view size is not compatible with input tensor's size and stride … Use .reshape(...) instead.`

**If you are stuck: My model runs, but some blocks never change during training.**

Check how you stored the blocks. A model is a class: `__init__` builds the parts, and `forward` uses them.
PyTorch finds a part’s weights only if the part is stored as a module. A plain Python list hides it:
`self.blocks = [Block(...), Block(...)]` runs fine, but `model.parameters()` doesn’t see those blocks, so the optimizer
never updates them. Use `nn.ModuleList([...])`. A single learned number needs the same care: wrap it in
`nn.Parameter(...)`, like the RMSNorm gain, or it isn’t trained either.

**Deeper: A GPT block is level 16’s decoder block, with RMSNorm**

Level 16’s decoder block was already pre-norm, and level 17 stacked it into a whole model.
Each block does `x = x + attn(norm1(x))`, then `x = x + ffn(norm2(x))`.
The attention uses a causal mask, and one final norm comes before the output layer.
A GPT block keeps all of it with one change: **RMSNorm** (level 19) instead of LayerNorm. Two more changes sit outside
the blocks: the positions are learned, and the token is not multiplied by √`d_model`.

The reference solution has 2 blocks, `d_model` 64, 4 heads, `d_ff` 256, learned position embeddings and RMSNorm:
102,415 parameters.

Now put the whole attention together. In NumPy, `k.transpose(0, 1, 3, 2)` swaps the last two axes of a four-axis array
(PyTorch: `k.transpose(-2, -1)`), and `np.where(blocked, -np.inf, scores)` sets the blocked scores to −∞ (PyTorch:
`scores.masked_fill(blocked, float('-inf'))`). The causal mask is the one from level 15: True above the diagonal.
Inside the function, take the sizes from the arrays themselves, for example `d_k = q.shape[-1]`.

**Code question.** Finish masked multi-head attention. q, k, v are already split into heads, shape (B, h, L, dₖ).

Fill in the blank (`____`):

```python
def causal(L):
    return np.triu(np.ones((L, L), dtype=bool), k=1)       # True = blocked

def split_heads(x, h):
    B, L, d = x.shape
    return x.reshape(B, L, h, d // h).transpose(0, 2, 1, 3)

def join_heads(y):
    B, h, L, d_k = y.shape
    return y.transpose(0, 2, 1, 3).reshape(B, L, h * d_k)

def attention(x, Wq, Wk, Wv, Wo, h):
    q, k, v = split_heads(x @ Wq, h), split_heads(x @ Wk, h), split_heads(x @ Wv, h)
    ____
    return join_heads(out) @ Wo

# 1 example, 3 tokens, d = 4
x = np.array([[[1.0, 0, 1, 0], [0, 1, 0, 1], [1, 1, 0, 0]]])
I = np.eye(4)
print(np.round(attention(x, I, I, I, I, 2), 4))
```

*Answer it on the page to check your work.*

When this works in NumPy, your PyTorch `Attention.forward` is the same four lines, translated.

**Boss check, part 1.** Your model is complete, but untrained. Run `python check.py my_gpt.py`. It builds your `GPT`
with small sizes, `GPT(d=32, heads=4, layers=2, d_ff=64)`, so keep these four argument names from the starter file
(`d` is `d_model`). Then it loads the fixed weights from `check_weights.json` into it by
name, runs it on `23+58=81` and turns 20 of its logits into a short code. It also tells you whether they match the reference.
A correct pre-norm GPT with RMSNorm, a causal mask and the √`d_k` scaling gives the reference’s code. A missing mask, a missing gain or a post-norm block each
change it. The weights are loaded by name, so one more thing must match: the order of the columns of `qkv(x)`. The first d
numbers are q, the next d are k and the last d are v (`q, k, v = self.qkv(x).split(d, dim=-1)`), and only then is each one split into heads.

**Code question.** python check.py `my_gpt.py` printed a line like forward = “0a1b2c3d4e5f”. Paste that whole line over the first line below, or put only the 12 letters and digits in place of ____.

Fill in the blank (`____`):

```python
forward = "____"      # 12 letters and digits, in quotes
print(forward)
```

*Answer it on the page to check your work.*

## Step 3: overfit one batch

Take 8 fixed examples and train on only those, for 300 steps. A correct model memorizes them:
with the loss on the answer digits only, the reference goes from 2.73 to 0.061 after 50 steps and 0.003 after 300.

The first number is not random. Before training, the model gives every one of the 15 tokens about the same probability,
1/15. The loss is −ln(1/15) = ln 15 ≈ 2.71. So an untrained model should start near 2.7.

**Predict.** The vocabulary has 15 tokens, and ln 15 ≈ 2.71. Your untrained model’s first loss on one batch is 9.4. What does that tell you?

A. Normal: a new model starts with a high loss
B. Something in the code is wrong: the output layer, the loss, or what the loss is counting
C. The model is too small

*Answer it on the page to check your work.*

If the loss on one batch stays high, the bug is in your code, not in the size of the model. It is almost always one of
three things: the causal mask, the shift by one, or what the loss is counting.

**If you are stuck: My loss on one batch stops at about 0.22 and won’t go lower.**

Check whether you count the loss on every position. The reference does exactly this: 0.36 after 50 steps,
then stays at 0.22. After `<bos>`, the 8 examples start with different digits, and no earlier token tells the model which digit comes first.
No model can get those guesses right, so the loss can’t go below that level.
Each example loses ln 8 ≈ 2.08 on its first digit. Spread over about 9.5 targets per example, that is ln 8 / 9.5 ≈ 0.22.

Count the loss on the answer digits only and the same model drops to 0.003.
Your model was fine. The test was asking it to guess a random digit.

**If you are stuck: The loss line fails with “Expected target size [8, 15], got [8, 10]”.**

`F.cross_entropy` wants one row of scores per prediction, but your logits are (8, 10, 15): 8 examples, 10 positions,
15 tokens. Flatten them to one row per position, as in level 17:
`F.cross_entropy(logits.reshape(-1, 15), y.reshape(-1), ignore_index=IGNORE)`. That is 80 rows of 15 scores and 80 targets.

## Step 4: train

Train on the 8,500 training problems, batches of 64, Adam with learning rate 0.001. One **epoch** is one pass over all
8,500, in a new random order each time. Train for 40 epochs, about a minute on a laptop. After every epoch,
measure accuracy on the 500 validation problems. Accuracy can drop for an epoch or two and then recover, so keep the
best model. Each time validation accuracy beats its best so far, save the model for the check. The starter file’s `train` shows
the one line that does it. If even the best epoch is below 95%, train for more epochs (for example 60). 

**Predict.** Loss on every position. After 8 epochs, accuracy on the 500 validation problems is 16%, and since epoch 4 it has only risen slowly, from 9% to 16%. The training loss is still falling slowly, from 1.34 to 1.27. What do you do?

A. Stop. The model is too small for this.
B. Lower the learning rate and restart.
C. Keep training.

*Answer it on the page to check your work.*

Two real runs of the reference model, one for each loss setting. Drag the slider or press play. The lab shows accuracy
on the 1,000 unseen problems, only to watch the runs: the best epoch is still picked on the 500 validation problems.

*[Interactive lab: Plateau — open the page to use it]*

**If you are stuck: Why does accuracy stay the same for several epochs, then jump?**

Accuracy only counts answers that are completely right. A model that gets two of three digits right scores zero.
During the flat part the model is learning pieces: the answer has two or three digits, and the last digit depends on
the last digits of a and b. The training loss shows this. With the loss on every position it falls from 1.34 to 1.27
between epochs 4 and 8, while accuracy only rises slowly, from 8% to 15%.

Then the model combines these pieces, including the carry. The carry is the 1 that moves to the next digit when a column
adds to 10 or more, as in 8 + 5 = 13. After that, whole answers start to be right.
How long the flat part lasts depends on the model, the data and the settings. In other setups it can last fifty epochs.

**Plateau or bug?** Look at the training loss, not only the accuracy.

- **The loss is still falling:** the model is learning. Keep training, even if accuracy hasn’t moved.
- **The loss has been flat for a long time and overfitting one batch (step 3) also fails:** it’s a bug. Train on
  something very easy, like the copy task from level 17, and fix the code before you train longer.

**If you are stuck: I ran it twice and got exactly the same numbers. Shouldn’t it be random?**

`torch.manual_seed(0)` fixes every random choice: the starting weights, the order of the batches and the split.
This is on purpose. When you change one thing in an experiment, the change you see comes from that one thing.
To see how much randomness matters, change the seed: the flat part gets longer or shorter, and the end result is similar.

**Predict.** With the loss on every position, the training loss stays around 1.0 forever, while the answers become almost all correct. Why can’t it go lower?

A. The model is too small to learn addition perfectly
B. The numbers before “=” are random, so no model can predict them
C. The learning rate is too high to settle

*Answer it on the page to check your work.*

**Deeper: Why the loss on the answer digits only starts faster**

With the loss on every position, most of the targets are the digits of a and b. They are random, so the model can’t
learn them, and they cause most of every batch’s gradient. With the loss on the answer digits only, all of the
gradient pushes the model toward adding correctly.

In these two runs, the run with the loss on the answer digits only is far ahead early: 81% after 8 epochs, against 15% for every position.
It reaches 95% a little sooner (epoch 14 against epoch 16). Both end near 100%. Neither curve is smooth: the run with
the loss on the answer digits only drops to 84.2% at epoch 18 and recovers. Look at several epochs before you decide a run is good or bad.

## Step 5: evaluate, and pass the boss

Give the model `<bos>a+b=` for each unseen problem, pick the top token, append it, and repeat until `<eos>`.
An answer counts only if every digit is right.

Don’t look only at the score. A model at 98% still gets 20 problems wrong, and the wrong ones often have something in common.
Print every miss and look for the pattern. `zip` reads two lists together, one pair at a time:
`list(zip([1, 2], ["a", "b"]))` is `[(1, "a"), (2, "b")]`.

**Code question.** Return every problem the model got wrong, written like ‘2+7=1’. problems is a list of (a, b) pairs, answers the model’s answers as strings, in the same order.

Fill in the blank (`____`):

```python
def misses(problems, answers):
    return ____

problems = [(2, 7), (50, 57), (4, 26), (17, 69)]
answers = ["1", "107", "20", "86"]
print(misses(problems, answers))
```

*Answer it on the page to check your work.*

**Boss check, part 2.** Run `python check.py my_gpt.py model.pt`. It first runs part 1 again on `my_gpt.py`, so the
model it tests must be the GPT you wrote. Then it loads your trained model and writes the answers to the
1,000 unseen problems with its own greedy loop. It prints a report with three parts: how many are right, the first ten
it got wrong, and a short check code made from those numbers. Paste the report.

The page can test that the report is consistent and that you pasted it unchanged. It can’t see your model: anyone who
reads check.py could type a report by hand. So this part relies on your honesty. The level counts as cleared when both parts pass.

**Code question.** Paste the report line that python check.py `my_gpt.py` model.pt printed: everything after “report =”.

Fill in the blank (`____`):

```python
report = ____      # for example {"correct": 961, "wrong": ["…", …], "check": "…"}
print(report["correct"], "of 1000 right")
```

*Answer it on the page to check your work.*

**Try it**

Once you pass, make it harder for yourself. Use 1 block instead of 2, or `d_model` 32 instead of 64.
Does it still reach 95%? How many epochs does it need now? Then replace the learned position embeddings with RoPE
(level 19) and compare how fast the two versions learn.

**Deeper: The reference solution**

The numbers on this page come from the course’s own solution to this boss. It is not a download, so that the boss stays a test.
Trained for 40 epochs (about a minute on a laptop), its best epoch on the 500 validation problems gets 98.6% of the 1,000 unseen problems right. The two runs in the lab above come from the same code. It has no learning-rate schedule, so their first 40 epochs are exactly a 40-epoch run; the lab shows 60 epochs, to show what happens after that.

You have now built and trained a GPT. If you haven’t taken them yet, two side trips show more.
Side trip [N7](/learn/encoder-decoder/) adds two things to the GPT you just wrote: a second stack, the encoder,
and cross-attention, which joins the two stacks. Side trip [U6](/learn/attention-backward/) computes the backward pass of attention by hand.

## You can now

- Turn a problem into ids, then into x and y shifted by one, with the loss on the answer digits only.
- Write masked multi-head attention: split the heads, mask the scores, merge the heads back.
- Debug a model before training it for long: check the first loss, then overfit one batch.
