# 1. Matrices and shapes

> How do you multiply two tables of numbers, and what shape is the result?

LLM by Hand · Foundations · runs in your browser · interactive page: https://llm.liko.page/learn/matrices/

Look inside a Transformer and almost every piece is a matrix multiplication. Practice this one operation
until it feels easy. After that, the rest of the course is much easier.

A **matrix** is a table of numbers. Its **shape** says how big the table is: shape (2, 3) means 2 rows and 3 columns.
`C[0][1]` means the number in row 0, column 1 of the matrix `C`.

## 1. One cell at a time

There is one rule: **the cell in row i, column j of the result is row i of the left matrix · column j of the right matrix.**
The “·” (the **dot product**) means: multiply the two lists number by number, then add everything up. Rows and columns count from 0,
so "row 0" is the top row and "column 1" is the second column, read top to bottom. Point at a cell of C (or tap it on a phone) to see it.

*[Interactive lab: Matmul — open the page to use it]*

**Question.** A = [[1, 2, 3], [0, 1, 0]] and B = [[1, 0], [0, 1], [1, 1]] What is C[0][1]?

*Answer it on the page to check your work.*

## 2. The shape rule

The shape rule comes from that one rule. To multiply a row of A with a column of B number by number,
the row and the column must have the same length. In NumPy, and everywhere in this course, `@` means matrix multiplication. So:

$$
(n, k) \; @ \; (k, m) \;\to\; (n, m)
$$

The two inner numbers must match, and they disappear. The two outer numbers are what is left. When the inner numbers match, we say the shapes **line up**.

**Question.** What is the shape of (3, 4) @ (4, 2)?

*Answer it on the page to check your work.*

**Predict.** A is (3, 2) and B is (3, 2). What happens with A @ B?

A. You get a (3, 2) result
B. You get a (3, 3) result
C. Error: the shapes don’t match

*Answer it on the page to check your work.*

The flip in that answer has a name: the **transpose**, written `B.T`. Row i of `B` becomes column i of `B.T`,
so a (3, 2) matrix turns into (2, 3). You will use it in level 6 and again in attention.

**If you are stuck: Why not just multiply cell by cell, like adding two matrices?**

That operation exists too. In NumPy it is `A * B`. It needs both matrices to have the same shape,
and it is used in places like masks. But it never mixes information between positions:
cell (0,0) only ever sees cell (0,0).

Matrix multiplication does mix. Every output cell combines a whole row with a whole column.
That mixing is exactly what a neural network layer needs. Each output is a weighted sum of all the inputs.

## 3. Write it yourself

Write the multiplication with three loops, without `@` or `np.dot`. The innermost loop runs over the inner size,
which is the dimension that disappears.

**Code question.** Write matrix multiplication with three loops, without @ or np.dot (a and b are lists of lists). Loop over the rows of the output, its columns, and the inner size, and add one product to an output cell on each pass. Several lines.

Fill in the blank (`____`):

```python
def mm(a, b):
    n, k, m = len(a), len(b), len(b[0])   # a is (n, k), b is (k, m)
    out = [[0] * m for _ in range(n)]      # (n, m), all zeros
    ____
    return out

print(mm([[1, 2, 3], [0, 1, 0]], [[1, 0], [0, 1], [1, 1]]))
```

*Answer it on the page to check your work.*

## 4. Shapes you will see a hundred times

A real model does not multiply one sentence at a time. It stacks several sentences into a **batch**, so the input
has three dimensions: (batch, length, features). Multiplying by a weight matrix only touches the last dimension.
The rule is the same: the inner numbers must match and disappear.

**Question.** X is (2, 5, 4): 2 sentences, 5 words each, 4 numbers per word. W is (4, 8). What shape is X @ W?

*Answer it on the page to check your work.*

Here is that multiply as blocks. Drag to turn it: the batch axis points away from you.

*[Interactive lab: Batch tensors — open the page to use it]*

**Deeper: Why a layer is (batch, length, `d_in`) @ (`d_in`, `d_out`)**

Think of the left matrix as "one row per word". Each word is a list of `d_in` numbers.
Think of the right matrix as "one column per new feature". Each column says how much of each input number
goes into that new feature.

So `X @ W` turns every word's `d_in` numbers into `d_out` new numbers, and every word uses the same `W`.
Nothing about the batch or the sentence length is in `W`. That is why the same layer works for
a batch of 2 or 2000, and for a sentence of 5 words or 500.

In level 14 you will see `X @ W_Q`, `X @ W_K`, `X @ W_V`: three of these, one after the other.

**Try it**

In the lab, add a column to A. The result disappears, because the shapes no longer match.
Now add a row to B. It comes back. Which number did you have to change, and why?

## 5. Broadcasting: adding arrays of different shapes

Adding two matrices of the same shape is easy: add cell by cell. But models constantly add things that are
*smaller*: one bias row to every example, one mask row to every word. NumPy has a rule for that, called
**broadcasting**: a dimension of size 1 is stretched (copied) until it matches the other array.

A column (4, 1) plus a row (1, 5): the column is copied across 5 times, the row is copied down 4 times,
and then they are added cell by cell.

**Question.** A column of shape (4, 1) plus a row of shape (1, 5). What shape is the result?

*Answer it on the page to check your work.*

That case had a 1 in each array. What if one array has no 1 at all?

**Predict.** A is (3, 2). v is a plain list of 3 numbers, shape (3,). Guess before you check: what does A + v do?

A. Adds v[i] to row i of A
B. Adds v to every row of A
C. Error

*Answer it on the page to check your work.*

Pick a case and point at or tap a cell of the result. The pale cells are the stretched copies. NumPy never stores them.
The lab uses its own numbers (a = [0, 1, 2, 3]), so its table is not the answer to the question below.

*[Interactive lab: Broadcast — open the page to use it]*

The rule, checked one dimension at a time **from the right**:

1. If the two sizes are equal, fine.
2. If one of them is 1, stretch it to the other.
3. Otherwise, error.

A missing dimension on the left counts as 1, so a (5,) array behaves like (1, 5).

**Question.** a = [[1], [2], [3], [4]] has shape (4, 1). b = [[10, 20, 30, 40, 50]] has shape (1, 5). What is (a + b)[2][3]?

*Answer it on the page to check your work.*

## 6. Where you will need it

This section looks ahead. Levels 10–15 explain its words; only the shapes matter here.
In levels 14–15 the attention scores of a batch have shape (batch, words, words): for every sentence,
every word's score for every other word. Some sentences are shorter and end in padding words, so each sentence
comes with a padding mask: `True` marks a padding word that must be blocked (its score is set to −∞ before
softmax, which you will meet in level 10). The mask is stored once per sentence, as (batch, 1, words),
and broadcasting stretches it to every row.

**Question.** Only the shapes matter here. Attention scores for a batch have shape (2, 4, 4): 2 sentences, 4 words each. The padding mask has shape (2, 1, 4). What shape do they broadcast to when the mask is applied to the scores?

*Answer it on the page to check your work.*

The same thing in 3D. Each sentence stores one mask row; broadcasting reads it for all 4 rows.

*[Interactive lab: Mask tensors — open the page to use it]*

**If you are stuck: Does NumPy really copy the column 5 times? Isn’t that wasteful?**

No copy is made. NumPy remembers that the dimension has size 1. Each time a cell of the stretched array is read,
it reads the same stored number again. That is why those cells in the lab are pale: they are not real cells.
So broadcasting a (1, 5) row against a million rows costs no extra memory for the row.

**Deeper: The bias in X @ W + b is a broadcast too**

A layer computes `X @ W + b`. `X @ W` is (N, `d_out`): one row per example. `b` is (`d_out`,), one number per
output. By the rule, `b` is treated as (1, `d_out`) and stretched down to N rows, so every example gets the
same bias. Nobody writes the copy; broadcasting does it. You will use it in every layer, starting in level 5.

## 7. Adding up along one axis

Many steps add numbers up along one direction of a table. In NumPy you choose the direction with `axis`.
Take A with shape (3, 2):

```
A = [[1, 2],
     [3, 4],
     [5, 6]]
A.sum(axis=0)  →  [9, 12]       # go down the rows: one total per column, shape (2,)
A.sum(axis=1)  →  [3, 7, 11]    # go across a row: one total per row, shape (3,)
A.mean(axis=0) →  [3, 4]        # the same, divided by the 3 rows
```

`axis=0` removes the rows. `axis=1` removes the columns. The axis you name is the one that disappears.

**Question.** X has shape (4, 3). What is the shape of X.sum(axis=0)?

*Answer it on the page to check your work.*

Now put the two ideas together. `X.mean(axis=0)` has shape (d,): one mean per column. Subtracting it from `X`
(shape (N, d)) is a broadcast: the row of means is stretched down to all N rows. The result has every column
centered at 0. LayerNorm in level 16 centers in the same way, but each row instead of each column: it takes
the mean along the last axis, one mean per word, and subtracts that.

**Code question.** Subtract each column’s mean from X, so every column has mean 0. Use .mean(axis=0) and broadcasting.

Fill in the blank (`____`):

```python
def center(X):
    return ____

X = np.array([[1., 2.], [3., 4.], [5., 9.]])
print("column means:", X.mean(axis=0))
print(center(X))
```

*Answer it on the page to check your work.*

**Try it: Read a line of data code**

Demo files often build a small dataset in one or two lines. Read them shape by shape:

```
# stack 60 rows on top of 60 rows → (120, 2)
pts = np.vstack([np.zeros((60, 2)), np.ones((60, 2))])
# 60 zeros, then 60 ones → (120,)
labels = np.repeat([0, 1], 60)
```

`np.vstack` puts arrays on top of each other, so the row counts add up: 60 + 60 = 120. `np.repeat` repeats each
item. When a demo line confuses you, print `.shape` after each step.

Optional side trip: [level U1](/learn/tensors-in-memory/) shows how the computer stores an array like these as one
row of numbers, and why some reshapes cost nothing. Level 2 continues the main line; you can come back to U1 any time.

## You can now

- Compute any cell of `A @ B` by hand: row i of A times column j of B, summed.
- Say the output shape of a matrix product, or that it fails, before running it.
- Predict the shape NumPy broadcasts two arrays to, and reduce along one axis.
