# N1. Convolutions

> How can one small filter find an edge anywhere in a picture?

LLM by Hand · Foundations · side trip: Classic networks · runs in your browser · interactive page: https://llm.liko.page/learn/convolutions/

A 28×28 picture is 784 numbers. An MLP from level 5 gives every one of them its own weight to every hidden unit.
But an edge looks the same whether it is in the top-left corner or the middle.
A **convolution** uses one small grid of weights, the **filter**, and slides it over the whole picture.

## 1. One filter, one window

The picture below is 5×5: a vertical stroke of 1s, like the digit 1. The filter is 3×3.
To get one output number, lay the filter on a 3×3 window of the picture, multiply the 9 pairs, and add them up.
Then slide one step and do it again.

*[Interactive lab: Conv — open the page to use it]*

**Question.** The picture is 5×5 with 1s in column 2 and 0s everywhere else. The filter is [[−1, 0, 1], [−1, 0, 1], [−1, 0, 1]]. What is out[0][0], the output for the top-left 3×3 window?

*Answer it on the page to check your work.*

The left side of the filter is −1 and the right side is +1. So the output is large where the picture is brighter on the
right than on the left. That is a vertical edge. Each output cell answers one question:
“is there a dark-to-bright step here, going left to right?”

**Question.** The picture is 5×5 with 1s in column 1 and 0s everywhere else. The filter is [[−1, 0, 1], [−1, 0, 1], [−1, 0, 1]], no padding, stride 1. What is out[0][1]?

*Answer it on the page to check your work.*

**Predict.** A 5×5 picture has 1s in column 2 and 0s everywhere else. With the horizontal-edge filter [[−1, −1, −1], [0, 0, 0], [1, 1, 1]], no padding and stride 1, what does the 3×3 output look like?

A. The same as before: 3, 0, −3 in every row
B. All zeros
C. 3, 0, −3 down every column

*Answer it on the page to check your work.*

**If you are stuck: Why does one stroke appear twice, as +3 and as −3?**

The stroke has two edges. On its left side the picture goes from 0 to 1: dark to bright, so the filter says +3.
On its right side it goes from 1 back to 0: bright to dark, so the filter says −3.
A network learns to use both. Later layers can combine “a +3 here and a −3 two pixels to the right” into “a thin line”.

## 2. Output size: stride and padding

Two settings change the output size:

- **Stride** is how far the filter moves each step. Stride 2 uses only every second position (0, 2, 4, …) and roughly halves the output.
- **Padding** adds a ring of zeros around the picture, so the filter can also sit on the border pixels.

For a picture of side $n$, a filter of side $k$, stride $s$ and padding $p$:

$$
\text{output side} = \left\lfloor \frac{n - k + 2p}{s} \right\rfloor + 1
$$

The brackets $\lfloor \; \rfloor$ mean **round down** to a whole number: $\lfloor 15.5 \rfloor = 15$.

Try both settings in the lab above. With padding 1 and stride 1, a 5×5 picture stays 5×5.

**Question.** A 28×28 picture, a 5×5 filter, no padding, stride 1. How many pixels wide is the output?

*Answer it on the page to check your work.*

**Question.** A 32×32 picture, a 3×3 filter, padding 1, stride 2. How many pixels wide is the output?

*Answer it on the page to check your work.*

## 3. Why one shared filter

The same 9 weights are used at every position. That is called **weight sharing**. Compare it with a dense layer
(level 5) that produces the same output: it has one weight for every pair of (input pixel, output cell).

**Question.** A 6×6 picture and a 3×3 filter give a 4×4 output. A dense layer that makes the same 4×4 output connects every one of the 36 pixels to every one of the 16 output cells. How many weights does it have (no biases)?

*Answer it on the page to check your work.*

For a real digit the gap is huge. A 28×28 picture has 784 pixels, a 26×26 output has 676 cells, and a dense layer needs
784 × 676 = 529,984 weights. The filter still needs 9: about 59,000 times fewer.
The filter learns “what an edge looks like” once and finds it everywhere.

**Predict.** Two networks train on 60,000 handwritten digits for the same 1,500 steps. A convolutional network has 5,258 weights. An MLP has 50,890. Which scores higher on 10,000 digits it has never seen?

A. The MLP: ten times more weights
B. The convolutional network
C. About the same

*Answer it on the page to check your work.*

**Deeper: Does a convolutional network still work on a shifted picture?**

Not fully. Moving the picture moves the feature map by the same amount: the filters find the same edges, just elsewhere.
But at the end, the network flattens the feature maps into one list and multiplies it by a dense layer. That layer
learned which positions matter, and it does not treat every position the same.

The level’s [demo.py](/files/convolutions/demo.py) measures it. Moving every test digit 3 pixels right and 3 pixels down drops the convolutional network from 98.0% to 31.2%,
and the MLP from 96.2% to 16.2%. Both are hurt. The convolutional network is hurt less, because its pooling
layers (section 5) keep only the largest number in each small block, so a small move changes less.

## 4. Channels: many filters at once

One filter answers one question, such as “is there a vertical edge here?”. A real layer asks many questions at once:
it has many filters, and each one makes its own output map. The maps are stacked, and each map is called a **channel**.
A layer with 8 filters turns one picture into 8 channels.

The next layer then reads all 8 channels. So each of its filters is not 3×3 but 8×3×3: one 3×3 slice per input channel.
At each position it multiplies all 8 windows by their slices and adds the 72 products into **one** number.
A color picture works the same way from the start: it has 3 input channels (red, green, blue).

| | shape |
|---|---|
| input | (`C_in`, H, W) |
| one filter | (`C_in`, k, k) |
| the layer | `C_out` filters, plus one bias each |
| output | (`C_out`, H′, W′) |

So a layer has $k \cdot k \cdot C_{in} \cdot C_{out}$ weights plus $C_{out}$ biases.
Example: a 3×3 layer from 1 channel to 8 channels has 3 · 3 · 1 · 8 = 72 weights and 8 biases, 80 numbers in all.

**Question.** A 3×3 convolution layer reads 8 input channels and makes 16 output channels. How many numbers does it learn, weights plus biases?

*Answer it on the page to check your work.*

That is how the convolutional network from the Predict above gets its 5,258 numbers: 80 for the first layer, 1,168 for the second,
and 4,010 for the final dense layer (16 channels × 5 × 5 positions = 400 inputs, to 10 digit scores: 400 · 10 + 10).

**Deeper: Same four steps, different connections**

An MLP, a convolutional network, a recurrent network (side trip N3) and a Transformer (level 14 onward) all do the same
four things: multiply inputs by weights, add a bias, apply a nonlinearity, and learn the weights with gradient descent.
They differ only in **which inputs each weight touches** and **which weights are shared**.

| network | each output reads | weights shared across |
|---|---|---|
| MLP | every input | nothing |
| convolution | a small window | every position in the picture |
| recurrent | this step’s input and the last state | every time step |
| Transformer | every token, mixed by attention weights | every token position |

In deep learning, “convolution” does not flip the filter first, as the math definition does. It is really a
*cross-correlation*. Because the filter is learned, the difference does not matter.

## 5. Stacking layers: seeing more of the picture

One 3×3 filter sees a 3×3 window. Put a second 3×3 layer on top of the first one. Each of its outputs reads a 3×3 window of
first-layer outputs, and each of those reads 3×3 pixels. Together, one second-layer output depends on a bigger patch of the picture.
That patch is its **receptive field**.

**Question.** Two 3×3 layers are stacked, stride 1, no pooling. One output pixel of the second layer depends on a square patch of the picture. How many pixels wide is that patch?

*Answer it on the page to check your work.*

**Pooling** shrinks a feature map. 2×2 max pooling keeps the largest number in each 2×2 block, so the map halves in each direction.

**Question.** 2×2 max pooling of [[1, 3, 0, 2], [4, 2, 1, 1], [0, 0, 5, 6], [1, 2, 7, 0]] gives a 2×2 result. What is its bottom-right number?

*Answer it on the page to check your work.*

To find the receptive field of a deeper stack, track two numbers as you go up, one layer at a time:

- **width**: how many pixels wide the patch is. Start at 1 (one pixel).
- **gap**: how many pixels apart two neighboring cells of the current map are. Start at 1.

Then apply one rule per layer:

- a k×k convolution adds (k − 1) × gap to the width;
- a 2×2 pooling adds 1 × gap to the width, and then doubles the gap (its output cells sit twice as far apart).

Worked example, the first two layers of “conv, pool, …”:

| layer | width | gap |
|---|---|---|
| start | 1 | 1 |
| 3×3 conv | 1 + 2 × 1 = 3 | 1 |
| 2×2 pool | 3 + 1 × 1 = 4 | 2 |

Without pooling the gap stays 1, which is why two 3×3 layers gave 5. Now continue the table for two more layers.

**Question.** Track the patch width and the gap layer by layer. A k×k conv adds (k − 1) × gap to the width. A 2×2 pool adds 1 × gap to the width, then doubles the gap. After a 3×3 conv and a 2×2 pool, the width is 4 and the gap is 2. Add a second 3×3 conv, then a second 2×2 pool. How many pixels wide is the patch one final output sees?

*Answer it on the page to check your work.*

## 6. What deeper layers see

*[Interactive lab: Receptive — open the page to use it]*

Early layers see small patches and find edges. Deeper layers see large patches and combine edges into loops, corners and whole digits.

Here is the whole digit network from section 4 in 3D. Each layer is a block: width and height are the picture, depth is the channels.

*[Interactive lab: Arch — open the page to use it]*

**Try it**

Switch the convolution lab to **Real digits** and try the vertical-edge filter on the 1 and on the 0. Then edit the filter so it gives large outputs
only on strokes that lean like “/”. Which three cells did you make positive?

## 7. Write it yourself

First, how to cut out a window in NumPy. For a 2D array, `x[0:3, 0:3]` takes rows 0–2 **and** columns 0–2 in one step: a 3×3 block.
The list habit `x[0:3][0:3]` does something else: `x[0:3]` takes rows 0–2, and the second `[0:3]` takes rows 0–2 **of that again**,
so you get all columns. For a window starting at row i and column j, write `x[i:i+3, j:j+3]`.

`*` multiplies two same-shape arrays cell by cell, and `.sum()` adds up every number in an array.

Write a convolution with two loops. For each output position, take the window under the filter and add up window × filter.

**Code question.** Write one line: the output at (i, j) is the 3×3 window starting at (i, j), times the filter, summed.

Fill in the blank (`____`):

```python
def conv2d(x, k):
    oh, ow = x.shape[0] - 2, x.shape[1] - 2
    out = np.zeros((oh, ow))
    for i in range(oh):
        for j in range(ow):
            out[i, j] = ____
    return out

img = np.zeros((5, 5)); img[:, 2] = 1
k = np.array([[-1, 0, 1], [-1, 0, 1], [-1, 0, 1]])
print(conv2d(img, k))
```

*Answer it on the page to check your work.*

Now let the filter side and the stride change. With filter side `f` and stride `s`, output cell (i, j) reads the window that starts at
row `i * s` and column `j * s`, so the window is `x[i*s:i*s+f, j*s:j*s+f]`. With no padding, the output side is the formula from section 2
with p = 0. In Python, `a // b` divides and rounds down: `7 // 2` is 3.

**Code question.** Write a convolution for any square filter and any stride, no padding. Set the output side m, make out, and fill every out[i, j]. Replace ____ with as many lines as you need.

Fill in the blank (`____`):

```python
def conv2d(x, k, s):
    n, f = x.shape[0], k.shape[0]   # picture side n, filter side f
    ____
    return out

img = np.arange(25.0).reshape(5, 5)
print(conv2d(img, np.ones((3, 3)), 2))
```

*Answer it on the page to check your work.*

## You can now

- Compute one output cell by hand: lay the filter on its window, multiply cell by cell, add the products.
- Compute a layer’s output size and its number of weights and biases before you build it.
- Write a convolution with any filter size and stride in NumPy, with two loops and a window slice.
