LLM by Hand

Math you need, on one page

This is all the math the course uses. Each idea has one small example with real numbers. Read it once now, or come back when a level sends you here. Each section ends with a short “Try it”: compute the answer on paper first, then open it to check.

e, exp and ln

e≈2.72e \approx 2.72 is a fixed number, like π\pi. exe^x (in code np.exp(x)) is “e multiplied by itself x times”. It works for any x, not only whole numbers, and it is always positive.

x−2−100.512
exe^x0.140.3711.652.727.39

“Multiplied by itself x times” only makes sense for whole numbers. For x = 0, x = 0.5 or a negative x, just read the table: e0=1e^0 = 1 and e0.5≈1.65e^{0.5} \approx 1.65.

A negative power means “one over”: e−1=1/e≈0.37e^{-1} = 1/e \approx 0.37. And e0=1e^0 = 1.

ln (the natural logarithm, np.log in code) undoes exe^x: it answers “e to what power gives this number?”. So ln⁡1=0\ln 1 = 0, ln⁡e=1\ln e = 1, ln⁡2≈0.69\ln 2 \approx 0.69 and ln⁡10≈2.30\ln 10 \approx 2.30. ln only takes positive numbers, and ln of a number below 1 is negative.

Four rules cover everything in the course:

ea+b=ea⋅ebln⁡(ab)=ln⁡a+ln⁡bln⁡(1/x)=−ln⁡xln⁡(ex)=xe^{a+b} = e^a \cdot e^b \qquad \ln(ab) = \ln a + \ln b \qquad \ln(1/x) = -\ln x \qquad \ln(e^x) = x
EXAMPLE−ln⁡0.1=−ln⁡(1/10)=ln⁡10≈2.30-\ln 0.1 = -\ln(1/10) = \ln 10 \approx 2.30. And ln⁡(2×5)=ln⁡2+ln⁡5\ln(2 \times 5) = \ln 2 + \ln 5: ln turns a product into a sum. That is why a loss over many words can add up one −ln⁡p-\ln p per word.

Level 2 draws its 3D loss surface with height ln⁡(1+loss)\ln(1 + \text{loss}). Adding 1 keeps the number positive, and ln pulls big values down a lot: ln⁡(1+0)=0\ln(1 + 0) = 0, ln⁡(1+9)≈2.30\ln(1 + 9) \approx 2.30, ln⁡(1+99)≈4.61\ln(1 + 99) \approx 4.61. A bigger loss is still higher.

Try it: what is −ln⁡0.25-\ln 0.25? (Use ln⁡4≈1.39\ln 4 \approx 1.39.)

0.25=1/40.25 = 1/4, so −ln⁡0.25=−ln⁡(1/4)=ln⁡4≈1.39-\ln 0.25 = -\ln(1/4) = \ln 4 \approx 1.39.

Used in: level 3 (cross-entropy), level 10 (softmax, perplexity), level 11.

Powers and square roots

x2=x×xx^2 = x \times x and x3=x×x×xx^3 = x \times x \times x. A square is never negative: (−3)2=(−3)×(−3)=9(-3)^2 = (-3)\times(-3) = 9. The square root x\sqrt{x} (np.sqrt) undoes a square: 9=3\sqrt{9} = 3. It always means the positive root.

In Python a power is written **: x ** 2 is x2x^2, and on a NumPy array a ** 2 squares every element. Don't write x ^ 2: in Python ^ means something else, and on floats it gives an error.

EXAMPLE0.255=(1/4)5=1/10240.25^5 = (1/4)^5 = 1/1024: multiply 0.25 by itself five times. A small number raised to a power shrinks fast.
Try it: what is (-2) ** 2 + 3 ** 2, and what is its square root?

(−2) × (−2) = 4 and 3 × 3 = 9, so the sum is 13 and its square root is √13 ≈ 3.61.

Used in: level 2 (squared error), level 6 (the slope of tanh), level 7 (Adam).

Σ, the sum sign

∑i=0n−1xi\sum_{i=0}^{n-1} x_i means “add xix_i for every i from 0 to n − 1”. The small number under a letter is its position, counted from 0 as in code: x0x_0 is the first x. (Many math books count from 1 instead; the sum is the same.)

EXAMPLEWith x = [2, 5, 1]: ∑ixi=2+5+1=8\sum_i x_i = 2 + 5 + 1 = 8. In NumPy that is x.sum().
Try it: with x = [4, −1, 3], what is ∑ixi2\sum_i x_i^2?

16 + 1 + 9 = 26. Square first, then add.

Used in: level 1 (each cell of a matrix product is a sum), level 10 (softmax divides by a sum).

Mean, variance and standard deviation

The mean is the sum divided by how many numbers there are. The variance is the mean of the squared distances from the mean. The standard deviation (std) is the square root of the variance: a typical distance from the mean, in the same units as the numbers.

EXAMPLEx = [1, 3, 5, 7]. Mean = 16 / 4 = 4. Distances from 4: −3, −1, 1, 3. Squares: 9, 1, 1, 9. Variance = 20 / 4 = 5. std = 5≈2.24\sqrt{5} \approx 2.24.
Try it: what are the mean and the std of [2, 4, 6, 8]?

Mean = 20 / 4 = 5. Distances: −3, −1, 1, 3. Squares: 9, 1, 1, 9. Variance = 20 / 4 = 5, so std = √5 ≈ 2.24. (The same spacing as the example, so the same std.)

Used in: level 7 (starting weights), level 10 (Gaussian noise), level 16 and N2 (LayerNorm), D1.

The derivative: a slope

The derivative of a function says how fast its output changes when its input changes a tiny bit. It is the slope of its graph at that point. You can measure it with two numbers: change the input a little, divide the change in output by the change in input.

EXAMPLEf(x)=x2f(x) = x^2 at x = 3. Try x = 3.01: f(3.01)=9.0601f(3.01) = 9.0601. The output grew by 0.0601 when the input grew by 0.01, so the slope is about 0.0601/0.01≈60.0601 / 0.01 \approx 6. The rule “slope of x2x^2 is 2x2x” gives exactly 2×3=62 \times 3 = 6.

When a function has several inputs, the partial derivative ∂L/∂w\partial L / \partial w is the slope when only w moves and the other inputs stay fixed. All the partial derivatives together are the gradient.

For a function with one input, the derivative is often written f′(x)f'(x), read “f prime of x”. For f(x)=x2f(x) = x^2, f′(x)=2xf'(x) = 2x, so f′(3)=6f'(3) = 6. A layer’s f′f' in level 4 is the slope of its activation function.

Try it: f(x)=x2f(x) = x^2 at x = 5. What is f(5.01) − f(5), and so what is the slope?

f(5.01) = 25.1001, so the change is 0.1001. Divided by 0.01 that is about 10, and the rule gives 2 × 5 = 10.

Used in: level 2 onward. Training is “move each parameter against its slope”.

The chain rule

If y depends on u, and u depends on x, then the slope of y with respect to x is the two slopes multiplied:

dydx=dydu⋅dudx\frac{dy}{dx} = \frac{dy}{du} \cdot \frac{du}{dx}
EXAMPLEu = 3x and y = u². At x = 1: u = 3. The slope of y with respect to u is 2u = 6. The slope of u with respect to x is 3. So the slope of y with respect to x is 6 × 3 = 18. Check: y = (3x)² = 9x², whose slope 18x is 18 at x = 1.
Try it: u = 2x and y = u². What is the slope of y with respect to x at x = 3?

u = 6. dy/du = 2u = 12, du/dx = 2, so dy/dx = 12 × 2 = 24. Check: y = 4x², slope 8x = 24 at x = 3.

Used in: level 2 (where the gradient formula comes from), level 6 (backpropagation is this rule, over and over).

sin, cos and polynomials

sin (np.sin) is a wave: it goes up and down between −1 and 1 forever. sin⁡0=0\sin 0 = 0, and it repeats every 2π≈6.282\pi \approx 6.28. Level 8 uses y=sin⁡(πx)y = \sin(\pi x) as a smooth curve to learn: for x from −1 to 1 it goes down to −1, back through 0, up to 1 and back to 0.

x−1−0.500.51
sin⁡(πx)\sin(\pi x)0−1010

cos (np.cos) is the same wave, shifted by a quarter turn: cos⁡0=1\cos 0 = 1. Angles in sin and cos are in radians, not degrees. A full turn is 2π≈6.282\pi \approx 6.28 radians, and a quarter turn is π/2≈1.57\pi/2 \approx 1.57. So 1 radian is about 57°.

EXAMPLEcos⁡1≈0.54\cos 1 \approx 0.54 and sin⁡1≈0.84\sin 1 \approx 0.84. At a quarter turn, cos⁡(π/2)=0\cos(\pi/2) = 0 and sin⁡(π/2)=1\sin(\pi/2) = 1.
Try it: what are cos⁡π\cos \pi and sin⁡π\sin \pi? (π is half a turn.)

Half a turn points the other way: cos π = −1 and sin π = 0.

A polynomial is a sum of powers of x, each times a number called a coefficient: c0+c1x+c2x2+…c_0 + c_1 x + c_2 x^2 + \dots Its degree is the biggest power. A degree-1 polynomial is a straight line; a higher degree can bend more times.

EXAMPLE1+2x+3x21 + 2x + 3x^2 has degree 2 and coefficients 1, 2, 3. At x = 2 it is 1+4+12=171 + 4 + 12 = 17.
Try it: what is 2−x+x32 - x + x^3 at x = 2, and what is its degree?

2 − 2 + 8 = 8. The biggest power is 3, so the degree is 3.

Used in: level 8 (fitting a curve with too many parameters), level 16 (position encodings use sin and cos), level 19 (RoPE turns vectors by angles in radians).

Very small and very large numbers: 1e−5 and powers of ten

Computers write 1e-5 for 1×10−5=0.000011 \times 10^{-5} = 0.00001 and 3e4 for 3×104=300003 \times 10^{4} = 30000. The number after the “e” says how many places to move the decimal point: a minus sign moves it to the left (a small number), a plus or no sign moves it to the right (a big number). This “e” has nothing to do with e≈2.72e \approx 2.72.

EXAMPLE2e-11 = 0.00000000002. 1.5e3 = 1500.
Try it: write 1e-3 and 4e2 as plain numbers.

1e-3 = 0.001 (three places to the left). 4e2 = 400 (two places to the right).

To multiply powers of ten, add the exponents: 103×104=10710^3 \times 10^4 = 10^7. To divide them, subtract the exponents: 1014/1010=10410^{14} / 10^{10} = 10^4. With a number in front, divide that number separately.

EXAMPLE1014/(1.4×1010)=104/1.4≈7,14310^{14} / (1.4 \times 10^{10}) = 10^4 / 1.4 \approx 7{,}143. And 1.4×1010/1012=1.4×10−2=0.0141.4 \times 10^{10} / 10^{12} = 1.4 \times 10^{-2} = 0.014.
Try it: what is 2.8×1013/10142.8 \times 10^{13} / 10^{14}?

Subtract the exponents: 13 − 14 = −1. So it is 2.8 × 10⁻¹ = 0.28.

Used in: level 6 (checking a gradient with a tiny step), U2, level 20 (memory and speed of a large model).

Vectors, matrices and shapes

A vector is a list of numbers, like [2, 0]. A matrix is a table of numbers; its shape (rows, columns) says how big it is. Rows and columns count from 0. Matrix multiplication A @ B needs the inner sizes to match: (n,k) @ (k,m)→(n,m)(n, k)\,@\,(k, m) \to (n, m). In formulas, two matrices written next to each other with nothing between them, like XWX W, mean the same product X @ W. Level 1 teaches all of this from scratch.

EXAMPLE[1, 2] @ [[3], [4]] = 1×3 + 2×4 = 11. Shapes: (1, 2) @ (2, 1) → (1, 1).
Try it: what shape does (4, 3) @ (3, 2) give? And (4, 3) @ (2, 3)?

(4, 2): the inner sizes are both 3. The second is an error: the inner sizes 3 and 2 don't match.

Used in: level 1 and every level after it.