LLM by Hand

Build your own GPT, by hand.

Start with one dot product. End with a working model you wrote yourself.
Every step is small enough for pen and paper. Try the first one right here.

cat = [2, 0]
dog = [1, 1]
cat · dog =2×1 + 0×1 =
multiply the pairs, then add

Know the basics? Each level opens with “Skip it with a short test”: answer a few questions and the level counts as cleared. Or turn on reading mode (in ⚙ settings) to just read, or start at level 11.

The map

Foundations, then Theory. Side trips are optional (the main levels never need them) and sit next to the level they open after. Three more parts are planned.

a level boss: build it yourself ★★★stars you earned

Side trip: Under the hood
The machinery PyTorch hides from you: tensors in memory, numbers in a computer, your own autograd, feeding data to a model, debugging a model (boss), and the backward pass of attention. Each level opens after the main level where it becomes useful: U1–U5 during Foundations, U6 after level 17.
Side trip: Classic networks
The networks that came before the Transformer, and the steps that led to it: convolutions, residuals, recurrent networks, LSTMs, the first attention, autoencoders, and the 2017 encoder–decoder Transformer. Opens after level 10; the 2017 Transformer (N7) opens after level 17.
Side trip: Diffusion
Another way to generate: start from noise and clean it up. D1 and D2 open after Foundations; D3 builds on attention and Transformer blocks (levels 14–16).
FoundationsTheoryUnder the hoodside trip · opens after level 1U1. Tensors in memoryU1Tensors in memoryUnder the hoodside trip · opens after level 3U2. Numbers in a computerU2Numbers in a computerUnder the hoodside trip · opens after level 6U3. Build your own autogradU3Build your own autogradUnder the hoodside trip · opens after level 9U4. Feeding data to a modelU4Feeding dataU5. Debugging a modelU5DebuggingClassic networksside trip · opens after level 10N1. ConvolutionsN1ConvolutionsN2. Residuals and normalizationN2Residuals and normsN3. Recurrent networksN3Recurrent networksN4. LSTM and GRUN4LSTM and GRUN5. Seq2seq and the first attentionN5Seq2seq + attentionN6. Autoencoders and VAEsN6AutoencodersDiffusionside trip · opens after level 10D1. Adding and removing noiseD1Adding and removing noiseD2. Sampling and guidanceD2Sampling and guidanceDiffusionside trip · opens after level 16D3. Latents and DiTD3Latents and DiTClassic networksside trip · opens after level 17N7. The 2017 encoder–decoder TransformerN72017 TransformerUnder the hoodside trip · opens after level 17U6. Backward through attentionU6Attention, backward1. Matrices and shapes1Matrices and shapesHow do you multiply two tables of numbers,and what shape is the result?you are here2. Gradient descent2Gradient descentHow does a machine learn from its mistakes?you are here3. Neurons and classification loss3Neurons and lossHow does a model answer yes or no, and how do youmeasure a wrong answer?you are here4. Activation functions4Activation functionsWhy does a network need a non-linear step, andwhich one should it use?you are here5. Multilayer perceptron5Multilayer perceptronWhat do you gain by stacking layersof neurons?you are here6. Backpropagation6Backpropagationboss · build it yourselfHow does the gradient flow back through the network,layer by layer?you are here7. Optimization7OptimizationHow do you make training fast and stable on real data?you are here8. Generalization8GeneralizationHow do you train a network that also workson data it has never seen?you are here9. From NumPy to PyTorch9From NumPy to PyTorchWhat does PyTorch do for you, andwhat is still your job?you are here10. Probability and sampling10Probability and samplingHow does a model “roll the dice” to pick thenext word?you are here11. What is a language model?11Language modelsHow can predicting the next word turn into writing?you are here12. Tokens: from characters to BPE12Tokens and BPEWhere does text get cut into pieces, and why doesit matter?you are here13. Embeddings and similarity13Embeddings and similarityHow does a token id become a list ofnumbers that means something?you are here14. Self-attention14Self-attentionHow does one word “look at” theother words?you are here15. Masks and multi-head attention15Masks and headsHow do you stop a word from looking at some words, andlook in several ways at once?you are here16. The parts of a Transformer16Transformer partsWhat do positional encoding, the feed-forwardnetwork, residuals, and LayerNorm each do, and howdo they make a decoder block?you are here17. The whole Transformer17The whole TransformerHow does a model produce text one word at a time?you are here18. Generation strategies18Generation strategiesGiven the model’s scores, how doyou actually pick the next word?you are here19. Modern blocks19Modern blocksWhat do today’s models change inside each block,and why?you are here20. Inference cost: memory and speed20Inference costHow much memory and time does a model need to writeone token?you are here21. Write your own GPT21Write your own GPTboss · build it yourselfCan you write a working model from a starter file thathas the names but no code?you are hereEngineering14 levels · planned· The workload ledger· Accelerators· Kernels and FlashAttention· Talking between GPUsand 10 moreTraining large models5 levels · planned· Pretraining: data and scale· Training stability· Fine-tuning and LoRA· Instruction tuning, alignmentand 1 moreApplications7 levels · planned· Prompts and in-context learning· Structured output and tool calls· Retrieval· RAGand 3 more
  1. Foundations
  2. 1Matrices and shapesHow do you multiply two tables of numbers, and what shape is the result?
  3. Side trip: Under the hoodOpens after level 1The machinery PyTorch hides from you: tensors in memory, numbers in a computer, your own autograd, feeding data to a model, debugging a model (boss), and the backward pass of attention. Each level opens after the main level where it becomes useful: U1–U5 during Foundations, U6 after level 17.
    1. U1Tensors in memory
  4. 2Gradient descentHow does a machine learn from its mistakes?
  5. 3Neurons and lossHow does a model answer yes or no, and how do you measure a wrong answer?
  6. Side trip: Under the hoodOpens after level 3
    1. U2Numbers in a computer
  7. 4Activation functionsWhy does a network need a non-linear step, and which one should it use?
  8. 5Multilayer perceptronWhat do you gain by stacking layers of neurons?
  9. 6Backpropagationboss · build it yourselfHow does the gradient flow back through the network, layer by layer?
  10. Side trip: Under the hoodOpens after level 6
    1. U3Build your own autograd
  11. 7OptimizationHow do you make training fast and stable on real data?
  12. 8GeneralizationHow do you train a network that also works on data it has never seen?
  13. 9From NumPy to PyTorchWhat does PyTorch do for you, and what is still your job?
  14. Side trip: Under the hoodOpens after level 9
    1. U4Feeding data
    2. U5Debugging
  15. 10Probability and samplingHow does a model “roll the dice” to pick the next word?
  16. Side trip: Classic networksOpens after level 10The networks that came before the Transformer, and the steps that led to it: convolutions, residuals, recurrent networks, LSTMs, the first attention, autoencoders, and the 2017 encoder–decoder Transformer. Opens after level 10; the 2017 Transformer (N7) opens after level 17.
    1. N1Convolutions
    2. N2Residuals and norms
    3. N3Recurrent networks
    4. N4LSTM and GRU
    5. N5Seq2seq + attention
    6. N6Autoencoders
  17. Side trip: DiffusionOpens after level 10Another way to generate: start from noise and clean it up. D1 and D2 open after Foundations; D3 builds on attention and Transformer blocks (levels 14–16).
    1. D1Adding and removing noise
    2. D2Sampling and guidance
  18. Theory
  19. 11Language modelsHow can predicting the next word turn into writing?
  20. 12Tokens and BPEWhere does text get cut into pieces, and why does it matter?
  21. 13Embeddings and similarityHow does a token id become a list of numbers that means something?
  22. 14Self-attentionHow does one word “look at” the other words?
  23. 15Masks and headsHow do you stop a word from looking at some words, and look in several ways at once?
  24. 16Transformer partsWhat do positional encoding, the feed-forward network, residuals, and LayerNorm each do, and how do they make a decoder block?
  25. Side trip: DiffusionOpens after level 16
    1. D3Latents and DiT
  26. 17The whole TransformerHow does a model produce text one word at a time?
  27. Side trip: Classic networksOpens after level 17
    1. N72017 Transformer
  28. Side trip: Under the hoodOpens after level 17
    1. U6Attention, backward
  29. 18Generation strategiesGiven the model’s scores, how do you actually pick the next word?
  30. 19Modern blocksWhat do today’s models change inside each block, and why?
  31. 20Inference costHow much memory and time does a model need to write one token?
  32. 21Write your own GPTboss · build it yourselfCan you write a working model from a starter file that has the names but no code?
    1. Planned
    2. Engineering · 14 levels, planned
    3. Training large models · 5 levels, planned
    4. Applications · 7 levels, planned